Cari Blog Ini

Tampilkan postingan dengan label big data. Tampilkan semua postingan
Tampilkan postingan dengan label big data. Tampilkan semua postingan

Kamis, 17 Oktober 2013

Suggestions for governments stepping into open data


I've been completing a survey for the Spatial Industries Business Association (SIBA) related to the Queensland Government's open data initiative, where one of the questions asked Can you list or describe any learnings that would be useful in Queensland?



I've provided a number of my thoughts on this topic, having closely observed open data initiatives by government over the last five years, and written periodically on the topic myself, such as:






To share the thoughts I placed in the survey more broadly - for any value they have for other jurisdictions - I've included them below:




  • Data released in unusable formats is less useful - it is important to mandate standards within government to define what is open data and how it should be released and educate broadly within agencies that collect and release data.

  • Need to transform end-to-end data process. Often data is unusable due to poor collection or collation methods or due to contractual terms which limit use. To ensure data can be released in an open format, the entire process may require reinvention.

  • Open data is a tool, not a solution and is only a starting point. Much data remains difficult to use, even when open, as communities and organisations don't have the skills to extract value from it. There needs to be an ongoing focus on demonstrating and facilitating how value can be derived from data, involving hack events, case studies and the integration of easy-to-use analysis tools into the data store to broaden the user pool and the economic and social value. Some consideration should be given to integrating the use and analysis of open data into school work within curriculum frameworks.

  • Data needs to be publicly organised in ways which make sense to its users, rather than to the government agencies releasing it. There is a tendency for governments to organise data like they organise their websites - into a hierarchy that reflects their organisational structures, rather than how users interact with government. Note that the 'behind the scenes' hierarchy can still reflect organisational bias, but the public hierarchy should work for the users over the contributors.

  • Provide methods for the community to improve and supplement the open data, not simply request it. There are many ways in which communities can add value to government data, through independent data sets and correcting erroneous information. This needs to be supported in a managed way.

  • Integrate local with state based data - aka include council and independent data into the data store, don't keep it state only. There's a lot of value in integrating datasets, however this can be difficult for non-programmers when last datasets are stored in different formats in different systems.

  • Mandate data champions in every agency, or via a centre of expertise, who are responsible for educating and supporting agency senior and line management to adapt their end-to-end data processes to favour and support open release.

  • Coordinate data efforts across jurisdictions (starting with states and working upwards), using the approach as a way to standardise on methods of data collection, analysis and reporting so that it becomes possible to compare open data apples with apples. Many data sets are far more valuable across jurisdictions and comparisons help both agencies and the public understand which approaches are working better and why - helping improve policy over time.

  • Legislate to prevent politicians or agencies withholding or delaying data releases due to fear of embarrassment. It is better to be embarrassed and improve outcomes than for it to come out later that government withheld data to protect itself while harming citizen interests - this does long-term damage to the reputation of governments and politicians.

  • Involve industry and the community from the beginning of the open data journey. This involves educating them on open data, what it is and the value it can create, as well as in an ongoing oversight role so they share ownership of the process and are more inclined to actively use data.

  • Maintain an active schedule of data release and activities. Open data sites can become graveyards of old data and declining use without constant injections of content to prompt re-engagement. Different data is valuable to different groups, so having a release schedule (publicly published if possible) provides opportunities to re-engage groups as data valuable to them is released.





Selasa, 03 September 2013

Weird and wonderful uses for open data - visualising 250 million protests and mapping electoral preferences


One of the interesting aspects about open data is how creatively it can be used to generate new insights, identify patterns and make information easier to absorb.



Yesterday I encountered two separate visualisations, designed on opposite sides of the world, which illustrated this creativity in very different ways.



First was the animated visualisation of 250 million protests across the world from 1979 to 2013 (see below).



Based on Global Database of Events, Language, and Tone (GDELT) data, John Beieler, a Penn State doctoral candidate, has created a visual feast that busts myths about the decline in physical protests as people move online and exposes the rising concerns people have around the world.



Imagine further encoding this data by protest topic and displaying trends of popular issues in different countries or states, or looking at the locations of protests in more detail to identify 'hot spots' - in fact John has done part of this work already, as can be read about in his blog (http://johnbeieler.org/)







Second is the splendid Senate preferences map for the 2013 Australian Government election, developed by Peter Neish from Melbourne.



Developed again from public information, this is the first time I have ever seen a map detailing the flow of preferences between political parties, and it illustrates some very interesting patterns.



The image below is of NSW Senate candidates, and thus is the most complex of the states, but shows how this type of information can be visualised in ways never before possible by citizens without the involvement of traditional media or large organisations.



For visualisations of all states and territories, visit Peter's site at http://peterneish.github.io/preferences/









These types of open data visualisation lend themselves to a change in the way the community communicates and offer both an opportunity and a threat to established interests.





Governments and other organisations who grasp the power of data visualisation will be able to cut through much of the chatter and complexity of data to communicate more clearly to the community, whereas agencies and companies who hang back, using complex text and tables, will increasingly find themselves gazumped by those able to present their stories in more visual and understandable forms.





We're beginning to see some government agencies make good use of visualisations and animation, I hope in the near future that more will consider using more than words to convey meaning.