Data Lives: Exhaustive Study Notes on Data Production and Society
The Proliferation and Evolution of Data in the Everyday Lexicon
Over the last decade, the term ‘data’ has transitioned from being the exclusive domain of researchers and professional administrators to a central part of the everyday lexicon.
Common terms now pervasive in society include: big data, open data, database, personal data, data-driven, data brokers, data analytics, data plan, data deluge, data activism, data science, data infrastructure, data justice, and the General Data Protection Regulation (GDPR).
The popular metaphor "data is the new oil" reflects its perceived value in the modern economy.
Before the current data revolution, data were viewed in common-sense terms as gathered facts, scientific measurements, or building blocks for information and knowledge.
Modern digital systems are often "black-boxed," meaning users have little understanding of their internal workings beyond the immediate interface and effects.
The Erosion of Trust and the Reality of Data Collection
Data-rich organizations, including private companies and state bodies, are secretive by design to protect commercial advantages or the integrity of their work.
Public trust in these systems has been undermined by significant scandals:
The Snowden revelations: Exposed large-scale government surveillance of citizens.
The Facebook/Cambridge Analytica scandal: Demonstrated how personal data profiles were used to target individuals with posts intended to influence voting preferences.
Frequent data breaches: Personal records, including sensitive identification and financial information, are accessed by unauthorized parties almost weekly.
Interaction with data-driven systems is often non-negotiable; providing data is the "price" paid for government services, entitlements, commercial platforms, and rewards programs.
Digital transactions (online shopping, credit cards, mobile apps, social media) make it nearly impossible to live an "analogue" life.
Critical Data Studies: From "Raw" to "Cooked" Data
Critical data studies is a field where academics, journalists, and civil rights advocates examine the nature and production of data and its transformative effect on life.
A central tenet is that data is never "raw"; instead, it is always "cooked" according to a specific recipe. This means data do not pre-exist their generation but are actively produced through specific procedures and instruments.
The life cycle of data includes generation, mutation (cleaning, wrangling, transforming, combining), and eventual deletion.
The production of data is influenced by:
Theories and concepts.
Research designs and protocols.
Standards and regulations.
Resources and finance.
Organizational processes and ethics reviews.
Political contexts.
Data systems are "socio-technical" systems, meaning they reflect human values, desires, and social relations as much as scientific principles.
Data Footprints, Shadows, and Commodities
Everyday life is saturated with digital devices like smartphones, self-tracking health devices, and smart home appliances (Alexa, Siri, smart TVs).
Data footprints are the records individuals choose to create.
Data shadows are records captured about individuals regardless of their intent.
Data are valuable commodities traded in a vast global market for purposes such as verification, profiling, and decision-making.
Data often "precede" individuals, determining outcomes for job applications, loan approvals, tenancies, or retail offers.
Methodology and Scope of "Data Lives"
The work adopts a reflexive standpoint, moving away from sterile, third-person academic narratives to include personal reminiscence, journalistic essays, and short stories.
It utilizes "recovered auto-ethnography," drawing on 30 years of experience in data work, infrastructure building, and policy advice.
Tactics of estrangement and defamiliarization (often used in Science Fiction) are employed to prompt critical reflection on society.
The geographical focus includes Ireland (both the North and the Republic), the United Kingdom, United States, China, Hong Kong, and Australia.
Academic and Professional Background of the Author
Doctoral work in the early 1990s involved testing the validity of methods used to measure geographic knowledge and the reliability of statistical analytics.
Significant projects led or contributed to:
The All-Island Research Observatory (AIRO): Compiled datasets across Ireland and Northern Ireland.
The Irish Qualitative Data Archive (IQDA): Stored and shared qualitative social science data.
The Digital Repository of Ireland (DRI): A trusted national repository for digital collections from libraries and museums.
Building City Dashboards (BCD): Created the Dublin and Cork Dashboards to track city performance.
The Programmable City: A five-year study on the role of software and data in urban governance.
Advisory roles include the Data Forum of the Department of Taoiseach, Irish Census Advisory Board, and the Audit Committee for the Irish Central Statistical Office.
Public policy work: Writing data-driven analyses for the blog "Ireland After NAMA" (2009–2015) concerning housing and population change.
Narrative Case Study: "Blind Data" — A Dialogue on Scientific Practice
Participants: Emma (an electronic engineer focusing on sound sensors) and Julian (an anthropologist studying technology and society).
Setting: A cafe meeting ahead of a wedding date.
Emma’s Perspective on Science:
Science is rational, logical, and objective.
Data (sound) exists as an electromagnetic frequency to be collected.
Values "mechanical objectivity" (minimizing bias and calibration errors).
Julian’s Perspective on Science:
Science is contingent and relational rather than essential and determined.
Instruments and choices made by researchers "cook" the data.
Data are "sketchy" and depend on the interpreter.
Technical Discussion of Sound Monitoring:
Sensors: Sonitus sensors.
Process: Measurements taken every minutes to save battery life and communication costs ( transmission).
Scales: Decibels () as a logarithmic scale developed by Bell Labs in , named after Alexander Graham Bell.
Formulaic context: A change in power by a factor of corresponds to a change in level.
Data processing: Use of a "smoothing algorithm" to control for "noise" in the data.
Julian's Thought Experiment: If a sound monitoring station operated for years, starting with manual recordings (a person looking at a needle on the hour) and ending with automated transmission, is the nature and quality of the data the same? Emma argues yes (due to control factors), while Julian argues the instrument and context change the nature of the data retrieved.
Themes and Challenges in the Data-Driven World
Spreadsheet Errors: A documented case in the Irish government where a spreadsheet error resulted in the loss of billion euros, unnoticed for over a year.
Data Interoperability: The difficulty of harmonizing data across different jurisdictions, using Metropolitan Boston and the Ireland/Northern Ireland border as case studies.
Smart Cities and Surveillance:
China: Use of mass surveillance and social credit scoring as a threat to democracy.
Smart City Testbeds: Implementation of data-driven management in residential areas.
Dataveillance in Australia: Unemployed volunteer firefighters losing benefits due to personal data shadows.
COVID-19 Data: How pandemic data reshaped daily life and directed intervention measures.
Civic Action: The role of civic hacking, citizen science, and data justice in challenging systemic institutional issues like racism.
Beyond how we understand what data are, and returning to the example of a tree falling in the woods, like Julian, I believe that capturing the sound and the trajectory of the fall as data is not straightforward. The process by which we create data, extract information and produce knowledge is constructed, not simply observed. They are shaped by education and convention, and mutate over time. Established wisdom on practising science is not the same now as it was in the past: new ideas, instruments, techniques, methods, approaches, concepts and theories emerge that change how we frame scholarship and measure, handle, analyze and interpret data.
Jim Gray13 argues that there have been four dominant scientific paradigms (generally agreed conventions and theories shared by scientists14) centred on epistemological approaches (ways of making sense of the world). The first of these he terms experimental science, which existed pre-Renaissance and tended to describe natural phenomena through an empiricist (weight of evidence) approach. The second was theoretical science, which was dominant throughout the Renaissance until the mid-20th century. Here, experimentation was used to test and derive theories in order to explain not just describe the world. Computational science, which seeks to simulate and model complex phenomena developed with the digital age. Here, the data employed might be synthetic, simulated and derived through modelling processes that extrapolate from base ‘raw’ data, for example in scenario building in weather and climate change forecasts, or predicting traffic patterns.
What this discussion reveals is that not only is data manufactured, but the approach to and process of manufacturing has changed over time. There are presently a number of different epistemologies employed in science in order to generate evidence and makes sense of the world.
Can you provide a rich definition of data?
Data is defined as a collection of facts, figures, or information collected for analysis. It represents the raw material from which knowledge is generated. Data can exist in various forms, such as numbers, words, measurements, or observations. In modern contexts, data are processed and analyzed using various methodologies, thus transforming into valuable insights that inform decision-making and enhance understanding of complex phenomena. Moreover, data are shaped by the processes employed in their collection, organization, and interpretation, reflecting the context in which they were generated and utilized.
What is Emma's understanding of data? How does Julian view data?
How has our understanding of data changed over time?
What are the four ways we make sense of the world around us?
How is data used?
What approach to understanding data do you find most convincing and why?