Skip to main content

Privacy, Secrecy, and Data: A Proposal Preserved With Its Flaws

·6865 words·33 mins
Lysander Demos
Author
Lysander Demos
Lysander writes at the intersection of individual rights and collective sovereignty. Drawing from the philosophical traditions of the Social Contract and the democratic spirit of the ancient Dēmos (the people), he advocates for a society where every person’s inherent rights are held sacred above the accumulated privileges of institutions, corporations, or political classes.

Foreword (2026)
#

What follows is a proposal I first published elsewhere in May 2020, reproduced here nearly whole — including the section where it visibly falls apart. I have resisted the urge to clean it up, because the falling-apart is the point.

The 2020 draft argued that privacy and secrecy corrupt data, that data disparity is the deep harm, and that the remedy was a public data lake: all the information, open to everyone, on equal terms. Equal access — equal power, one symmetry to rule them all. I tried to build it and failed.

I made honest efforts to conceptually build the lake, but symmetry kept breaking down into asymmetric access structures. The asymmetry was needed to smooth power disparity: medical records open to everyone mean your insurer and your employer read them too, so a credentialed tier appeared to protect the weaker party; hence asymmetry. I stopped writing at the point where the proposal section below trails off into open questions. The questions I attempted to address have waited six years for my return. The questions are real, the problems brought to light are real, the attempted solution was false.

My determination is that equal access was never the goal. The goal was, and remains, harder: a solution that offers the advantages of easily accessible data while preserving the protections the weak once received from privacy. This was a privacy we never questioned until our capacity to capture and record overwhelmed it. Until recently, walls, distance, and forgetting supplied privacy for free.

This republication is a construction log. My diagnosis largely survives — the privilege of information, the harm of disparity, the danger of official secrecy, and the injustice done to the uncounted. I have marked, in bracketed [2026] notes, each point where my thinking has evolved. The substantive rethink is sketched in the afterword. Where a passage stands without annotation, it stands.

Aside from those notes, I made minimal adaptations: described a social-media screenshot in text rather than showing it, replaced a chart with its citation, corrected obvious typos, removed one personal link, and stripped a URL that no longer resolves. The argument, including its errors, is as it was.

A note on reading this. The foreword and afterword carry my argument. The 2020 text between them builds the evidence. Read that evidence for full understanding or skim the bracketed [2026] notes to see where my thoughts developed. Either path arrives at the afterword.


The following was first published May 1, 2020.


In January of 2020, I posted the core of this idea to Facebook. It did not get the level of feedback I had hoped for.

I had hoped to get feedback on the idea to inform and clarify my thoughts. Without that, I have tried to expand the idea and try to cover not only the positives from the idea but also consider many of the negatives.

This is a draft document. I have placed it on the internet for reviewing purposes. Mainly it is incomplete.

Let’s first clarify some terms.

Data and information. I will use these two terms interchangeably even though there are some subtle differences. Data and information refer to things such as:

  • point measurements such as a time series of heart beats which may be averaged to a heart rate,
  • the arrangement of pixels in an image which includes color information,
  • a person’s name or date of birth,
  • reviews given to a movie, and
  • submarine plans stolen by a spy.

These are just a few examples of data and information.

Record. Data and information are often in record sets. For example: a medical record might contain a person’s name, their date of birth, smoking history, family history of heart disease, and their death from lung cancer.

Knowledge. Knowledge is a conclusion inferred or deduced from records. We often use knowledge as a shorthand for connecting data to decision-making. An example of knowledge is the statement that cigarette smoking causes cancer.

Quality. When discussing quality data, I mean data that is accurate, timely, and complete. Low quality data may be caused by low resolution in measurement, intentional skewing of data, poor communication of data, random errors in recording data, and delays in data reporting.

Privacy. This is a complicated term and one that will be discussed at length later. Generally, privacy is something that individuals but not organizations have. Privacy typically concerns the information of that individual. Often, privacy is considered a good thing and stated in terms of the right to be left alone.

Secrecy. A secret may be held by individuals and organizations. Often secrecy is considered a bad thing but there are often times when it is necessary and good. Coca-Cola’s formula is a trade secret which is not only allowable under the law but recognized as a proper way to conduct business. The invasion plans for D-Day were a secret and it should be clear that this is a legitimate use of secrecy.

Why data?
#

Data informs much of our decision making. If I know store A sells a gallon of milk for less than store B, I can save money. We also regularly use knowledge drawn from data. Knowing that smoking is often a cause of lung cancer, a person can choose to improve their health. If we know that a college degree leads to lower crime rates, as a society, we may choose to invest in better educational opportunities.

In the stock market, the price of a share of stock is determined by what a buyer is willing to pay and what a seller is willing to accept. Buyers and sellers are acting upon how they expect the company to perform in the future. So the pricing is determined based upon expectations of earnings, valuation, rates of return, market conditions, and emotions. Having accurate information about those items is necessary to properly price a share of stock.

With the introduction of Deep Learning methods in machine learning, the need for data has increased exponentially. These algorithms require big data to get good results. In many cases, more high quality data is the difference between an algorithm that barely performs as well as the average person and one that out-performs even the experts.

All of these are examples of ways in which having data may lead to better decisions for individuals and societies. Having data is not sufficient for making quality decisions but it is necessary. Those having more high quality data have the opportunity to improve their decision-making. Data is power especially if there exists a disparity with the information commonly known.

Data is power. With data we gain the power to effect change and advantage in our world.

We have unequal access to data
#

Using the stock market example, a person directly involved with the company will often know of dramatic changes in future earnings or big problems before the average investor. An insider acting on their information can make large profits. Insider trading on private information reduces the trust investors have in the company. If this reduction in trust becomes widespread, the markets no longer operate efficiently and profits accrue only to those with access to private information. Inefficient markets and lack of trust in the information provided by companies harms all. This has clearly been recognized by regulators and is why insider trading is a crime.

There are many other situations where data disparity is normal. Sometimes that disparity is related to having the skills to use the information, such as in the professions of medicine, law, and engineering. We rely on the skills but also on the ethical obligations of such professionals. Each of the mentioned professionals is expected to abide by standards and codes of ethics. Each of these professions also requires a license to practice and those licenses may be revoked for not following the ethics of providing fair and honest service. Again, we recognize that having an advantage in information is powerful and may be abused.

Holding information others do not have gives the holder easy-to-abuse power over others. Blackmail is an obvious example. Militaries seek advantage by having more complete information about their opponent than the opponent has on them. Many seek advantage by holding secret information. The problem is that secret information, especially when abused, leads to mistrust, questioning of motives for actions, and guessing at explanatory information.

Data disparity is harmful. Those without access to quality information can expect to make poorer decisions and lead less fulfilling, shorter, unhealthier, and poorer lives. This disparity also breaks down society, especially in trust relationships.

Let’s talk about quality
#

Quality data is accurate, timely, and complete. We rely on data to inform our decisions. Having quality data allows us to make better decisions faster and with a better understanding of all the things that might go wrong. Quality information is the new currency for opportunity.

Anything that degrades the quality of the data might be considered harmful. Sometimes data corruption is unavoidable. A heart rate monitor may be worn incorrectly or have a power failure. The clocks used to record separate but related data might not be synchronized.

It may be difficult to even know a problem exists if we do not record certain types of information. In 2016, I read The Vanishing of Canada’s First Nations Women. This article highlighted the problem of lack of data: Pearce enrolled in a doctoral program in law to research missing and murdered women but soon found that “there was nothing available to the public in terms of data” because police had never published national statistics.

Then, from the Urban Indian Health Institute report from 2018: “As demonstrated by the findings of this study, reasons for the lack of quality data include under reporting, racial misclassification, poor relationships between law enforcement and American Indian and Alaska Native communities, poor record-keeping protocols, institutional racism in the media, and a lack of substantive relationships between journalists and American Indian and Alaska Native communities.”

These articles highlight the need for quality data to determine if problems even exist. For data to have high quality, it must also be collected uniformly. Uneven data collection is a real problem especially when there are strong incentives to suppress correct reporting. Crime data is the obvious example of reporting discrepancies. Different jurisdictions report data differently and there are often incentives to under-report or reclassify certain types of crime. See Measurement Problems in Criminal Justice Research and a Journal Sentinel investigation that found the Milwaukee Police Department had underreported thousands of violent assaults, rapes, robberies and burglaries and failed to correct the problem while presenting flawed statistics to the public.

These articles also highlight the intersection of data power and data disparity.

Quality matters. Data is not enough; it must be accurate, timely, and complete to be truly useful. Poor quality data may be used as a lever of power.

Intentionally corrupt data
#

Certainly any type of intentional corruption of data should be unacceptable. Modifying data seriously harms the usefulness of the information derived. Furthermore, any decisions made based upon that data are likely to be wrong.

This might seem an unlikely problem but really it is everywhere. People regularly lie when filling out survey forms. In fact, many surveys have validation questions to correct and/or eliminate dishonest responses.

The Global Positioning System (GPS) was originally a Department of Defense project. When permitted for civilian use, the signal was intentionally degraded to prevent high accuracy. More recently, competition has forced the GPS signal to provide more accurate position.

Why would data be intentionally corrupted? There are many reasons but these include: to maintain an information advantage, to cause bad decision-making, and to maintain privacy.

Privacy often limits the amount of data collected, the timeliness of the data, and the accuracy of the data. Certainly, any types of anonymization techniques reduce the completeness of the data. Definitively, we can state that privacy reduces the quality of data and intentionally low quality data can cause harm.

Secrecy and privacy corrupt data and cause harm. We intentionally corrupt data to protect privacy and allow for secrets. We should try to minimize these corruptions.

[2026] Is privacy a primary cause of intentional data corruption? That was the concern I indicated, but I have since come to weigh the opposite force more heavily: observation itself corrupts data. Behavior changes when people know they are watched. Social science has recognized this since the Hawthorne studies, and names the general problem reactivity: the act of measurement alters the thing measured. A recent demonstration makes the scale concrete. When Bernstein and Turban (2018) tracked firms that removed spatial privacy by converting to open-plan offices, face-to-face interaction did not increase as intended — it fell by roughly seventy percent, as people replaced their lost private space with a manufactured privacy. A person who knows they are being watched performs. When the performative data is captured, it is already corrupt. Limited data due to privacy is better than corrupted data. More on this in the afterword.

The dark side of secrecy
#

In government
#

The following from Schoenfeld very eloquently states my general thoughts on secrecy in government:

“A basic principle of our political order, enshrined in the First Amendment guarantee of freedom of speech and of the press, is that openness is an essential prerequisite of self-governance. Indeed, at the very core of our democratic experiment lies the question of transparency. Secrecy was one of the cornerstones of monarchy, a building block of an unaccountable political system constructed in no small part on what King James the First had called the ‘mysteries of state.’ Secrecy was not merely functional, a requirement of an effective monarchy, but intrinsic to the mental scaffolding of autocratic rule.

Standing in diametrical opposition to that mental scaffolding was an elementary proposition of democratic theory: Legitimate power could rest only on the informed consent of the governed. Along with individuals at liberty to give or to withhold approval to their government, informed consent requires, above all else, information, freely available and freely exchanged. Official secrecy is anathema to this conception. No one has put this proposition more forcefully than James Madison, who tells us that ‘A popular government, without popular information, or the means of acquiring it, is but a Prologue to a Farce or a Tragedy, or, perhaps both. Knowledge will forever govern ignorance: And a people who mean to be their own Governors must arm themselves with the power which knowledge gives.’”

There are situations when secrecy is needed, most notably in cases of national security. Secrecy should be the exception and not the rule. It should require a clear statement of why something should be secret and then it should be made public as soon as the requirement for secrecy has passed.

Secrecy hides the decision-making process and consolidates power to those holding the secrets. It keeps the people uninformed, limits participation, and allows for corruption to take root and grow unchecked. We must, in order to remain in a free and functioning democracy, vigilantly limit secrecy in government at all times.

Governments have a preference for secrecy and the ability to act without the people’s oversight. Thus secrecy is a slowly encroaching action of government and must be constantly guarded against. Yes, it is increasing now, as Aftergood’s article from March 2020 states: “The Department of Defense is quietly asking Congress to rescind the requirement to produce an unclassified version of the Future Years Defense Program (FYDP) database.”

[2026] An interesting thing about secrecy is that every secret has an expiration, a time in which it ceases to exist as data or becomes known. The question is whether that timer runs by design, leak, or disinterest. Declassification schedules with expiring defaults should be the norm.

Limiting secrecy applies not only to the government but also the great influencers of government and to the tools used by governments. For influencers of government, I include things like: lobbyists, donations, political action committees, and those groups or individuals that gain influence using money or shared secrets. Finally, increasing citizen participation in local governance is an excellent way to both keep aware of encroaching secrecy and to also reduce it. See more at Global Answers for Local Problems, Lessons from Civically Engaged Cities.

In the business world
#

Honestly, this section needs more thought and development but here goes. Some ideas may be rather controversial because they feed into other not-fully developed thoughts I have on taxation policies. Another article in the future might address that issue but that is not as much in my core competencies as data is.

There are different types of businesses ranging from sole proprietorships to publicly traded corporations and the rules applying to them often differ greatly. I will limit discussion here to publicly traded corporations. Generally, there is already a lot of transparency in these businesses due to the required reportings to shareholders and government, but businesses can be complex, which gives opportunity for secrecy. Even with that transparency, there still exist many areas for improvement. The areas for improvement mainly cover influencing actions toward government and collection and handling of individuals’ data.

Lobbying should be fully disclosed. Sometimes, a business participates in lobbying to push forward legislation in an area where the business is an acknowledged expert. This is reasonable but their participation should be checked by participation for citizens or groups which might oppose the legislation. Lobbying that is of a political nature only should be prohibited. There are gray areas between purely expert and purely political and this is why their lobbying activities should be fully disclosed and scrutinized.

Charitable activities should be curtailed entirely as these are typically either marketing or lobbying activities in disguise. If executives of a business wish to be charitable, they should use their own funds to purchase the services of the business for donation and not impose their charitable preferences on their diverse shareholders.

[2026] The charitable activities paragraph is off-thesis; consider it retired.

Let’s now discuss the handling of individuals’ data. For some businesses, this is just a byproduct of interacting with customers but for others, this data is their lifeblood source of revenue. A company earning revenue based upon their database of individual users is not really paying for their access to raw material. They are also building barriers to entry to other companies based not upon their prowess or technical advantage but upon their access to the raw material.

The raw material is individual persons’ data, often data a person would consider private. The business considers this data as their property with the rights to sell, use, or keep it secret within lawful limits.

Unchecked secrecy corrupts. Clearly, the founders of the United States knew secrecy in government was dangerous. We now also see that secrecy in business allows for hidden influence and the co-opting of the people’s privacy and power.

Privacy’s offsetting benefits
#

A lot has been written in support of privacy and the right to privacy. In fact, until recently — driven by my interests in machine learning and my understanding of the harm caused by low quality data — I was a strong supporter of the right to privacy. I put both time and money into supporting privacy rights. So, let’s examine the reasons for privacy.

I’ll base this on Solove’s Conceptualizing Privacy and on Magi’s Fourteen Reasons Privacy Matters: A Multidisciplinary Review of Scholarly Literature, shown in the References section.

Solove identifies six general types of definitions of privacy:

  1. the right to be let alone,
  2. the ability to limit access to the self by others,
  3. secrecy or concealment of certain matters,
  4. the ability to control information about oneself,
  5. the protection of one’s personhood, individuality and dignity, and
  6. control over one’s intimate relationships or aspects of life.

The problem of corrupted data is mainly about information generated by a person or information about a person and not imposing upon or controlling the person.

Let’s look at this a bit deeper. Magi lists fourteen reasons. I’ll list them here and discuss a few of them in more depth for better understanding. These fourteen reasons will be addressed further in a later section.

  1. Privacy protects from overreach of social interactions and provides opportunity for relaxation and concentration.
  2. Privacy affirms self-ownership and the ability to be a moral agent.
  3. Privacy prevents intrinsic loss of freedom of choice.

These three reasons point to impositions on our private space to affect or direct our thoughts and ability to act.

  1. Privacy allows freedom from self-censorship and anticipatory conformity and allows people to explore their “rough draft” ideas.
  2. Privacy helps prevent sorting of people into categories that can lead to lost opportunities and deeper inequalities.
  3. Privacy prevents being misjudged out of context.
  4. Privacy provides a physical space in which an individual can control the artifacts that support the narrative of her/his life.
  5. Privacy preserves the chance to make a fresh start.
  6. Privacy allows individuals to be authentic and to play appropriate roles in various contexts.
  7. Privacy supports intimacy and the building of relationships.
  8. Privacy supports the common good.
  9. Privacy protects from power imbalance between individuals and government/organizations.
  10. Privacy supports democracy, political activity, and service.
  11. Privacy provides space in society for disagreement.

Privacy enhances people and society. The overall impression from this list is that privacy is both for the individual and for society. The benefits of privacy for the individual protect their physical, emotional, spiritual, and intellectual space. The benefits for society enhance innovation, equality, justice, involvement, and decrease conflict.

How much privacy do we have?
#

Over the past few years, many articles have lamented the erosion of personal privacy. Our every click may be monitored by our favorite website or social media company. With technological innovations, governments are able to track and monitor individuals at an unprecedented level. To prevent money flows to terrorist organizations, we have instituted rules and regulations to make financial transactions more traceable. The Health Insurance Portability and Accountability Act of 1996 (HIPAA) was enacted to protect the privacy of our health records. Cameras record our actions at intersections, walking down the street, and in both public and private spaces. Our current location is readily surrendered by the smartphone device we all carry. Every email, photo, and online interaction we engage in is recorded and saved for posterity.

So yes, your data, much of which you may consider private, is held by some faceless government, business, or organization. Do the faceless have your best interests in mind? I think not, and that is why I fought against this intrusion for many years.

It’s not really a matter of what information is out there but how consolidated and cohesive it is. The government has or can gain access to all of your data, and it may legally require that you never be informed. As this issues section from the Electronic Frontier Foundation points out, “The USA PATRIOT Act broadly expands law enforcement’s surveillance and investigative powers and represents one of the most significant threats to civil liberties, privacy, and democratic traditions in US history.”

Privacy is false anyway. Our data is out there. The question has become: who is using it?

[2026] This statement has evolved; by sharpening, not reversing. Privacy, as a fact about the world, is false: everything can be captured, and what once passed for privacy was only the expense of capture. That expense included walls, distance, darkness, forgetting. The expense of capture has been greatly reduced and I expect the cost reductions to accelerate. Once our data is captured, it is often permanently available. Who is using it becomes the question. The formation domain is where our political, intellectual, and personal life is built: the ballot, the library, association, worship, drafts, and thought. There, engineered non-capture can exist. Privacy in the formation domain is not a personal option. It is not a checkbox a company offers. We build zones where capture is prohibited by law and made impossible by design. The secret ballot is the model: a voter cannot prove their true vote even if they want to. The ballot holds because the booth is an engineered non-capture zone. The act leaves no record for anyone to demand. In this domain, we should presume that anything that can be disclosed can be demanded by anyone with leverage: an employer, a landlord, and a platform. Privacy that can be waived will be waived. In the formation domain, only what cannot be captured remains forever private. Elsewhere, data is captured and access is governed: medical, financial, and transactional. That is the accountable-institutional layer, not the formation domain.

The proposal
#

But the algorithm has overcome all.

Some years ago, when cameras were initially being installed in many public spaces and were being monitored by public officials, or more likely by algorithms, I saw a piece that suggested the only way to achieve détente was that viewing of the cameras should be equal access to all.

I propose the creation of a public data lake to hold this information. There could be different sub-module lakes that include financial or health data. Read access to the lake would be credentialed with some sort of credit/debit scheme. Changes and updates to the data would be through a pull request method. Initiation and oversight of this data lake would be done by some sort of government, citizen, and business consortium.

How is this possibly a good idea?

There is the problem of theft or use of information for nefarious purposes.

Share the database such that everyone has access to the data.

Problems with localism and fragility?

What to do here?

[2026] Clearly, I was struggling here to reach my selected solution. That proved to be a failure, recorded exactly as it stalled. The intent was a data lake with equal access for all, but that was not feasible. Every turn and complication added permission layers onto a supposedly simple solution. Why? Simply because the capability to act upon data was never dependent on access to the data alone. It had much more to do with the power disparity between individuals and other individuals, organizations, and governments. A corporation with analysts, lawyers, and compute reads the same lake very differently than a tenant does. Access to data still matters, but reining in privilege requires more than access: the encoded data must be made legible to the people. By legible I mean something beyond available. A nine-hundred-page regulatory filing is available; it is not legible. Data is legible when an ordinary citizen can extract what it means for them without hiring an expert — can see who benefits, who pays, who decided, and what authority is being claimed. Disclosure that only insiders can interpret is not transparency. The afterword takes this up.

Impacts on privacy
#

Let’s look at this a bit deeper. Magi lists fourteen reasons. So let’s address each in turn.

  1. Privacy protects from overreach of social interactions and provides opportunity for relaxation and concentration.
  2. Privacy affirms self-ownership and the ability to be a moral agent.
  3. Privacy prevents intrinsic loss of freedom of choice.

Quality data collection should not affect these three reasons. If we are speaking of the intrusion of unwanted people into social interactions, this may be a problem. Generally though, this is a problem that can be addressed by other legal means that might be supported by data. Stalking is an example of this. A stalker might try to inject themselves based on available information but their location might be legally used to prohibit and prosecute their actions.

[2026] This response was exactly backwards. See the afterword.

  1. Privacy allows freedom from self-censorship and anticipatory conformity and allows people to explore their “rough draft” ideas.

Without absolute privacy, people often engage in self-censorship and anticipatory conformity. Some self-censorship is beneficial but too much is harmful to society. Decreasing overall privacy will increase self-censorship, therefore we will need mechanisms to correct this imbalance. This imbalance may be somewhat offset by clear and strong laws to protect against official or societal curtailing of thoughts and ideas. We might also engage in positive reinforcement of diversity. Finally, the creation of strong anonymous channels may allow for the appropriate expression of ideas without oppression.

A question to consider is: are we able to measure how much is lost to self-censorship and conformity? If the loss is great and we are not able to mitigate that loss, that is a point upon which to reinstate strong data privacy.

[2026] The proposed fix — strong anonymous channels — is privacy reintroduced through the back door. And the closing question answers itself: the loss to self-censorship is large, and the reinstatement clause triggers. See the afterword.

  1. Privacy helps prevent sorting of people into categories that can lead to lost opportunities and deeper inequalities.

There may be some sorting of people into categories but at the same time opportunities will likely remain the same and inequalities should be lessened. In fact, the reduction of disparity and inequalities is one of the benefits of good data.

[2026] Inverted: open access makes sorting cheaper for every institution simultaneously. See the afterword.

  1. Privacy prevents being misjudged out of context.

Initially, a person’s data will be judged out of context. Having context to the data is generally an improvement such that the data will seek context. People with access to the data may not exercise the same discretion about including context with data. Perhaps this is an aspect that will take a little bit of time to find equilibrium.

  1. Privacy provides a physical space in which an individual can control the artifacts that support the narrative of her/his life.

An individual will not be able to control the digital artifacts in their life. False narratives will be very difficult to support. At the same time, true narratives will be easier to support and recall as the data is readily available to the person. If we talk only about physical spaces, then improved data should have little impact.

  1. Privacy preserves the chance to make a fresh start.

Higher quality data will likely make it more difficult to make a fresh start. We have already seen that just based on the longevity of static data. What was once forgotten is now stored. There may be some solutions for this which include legislation that rolls certain types of data into archival storage. Recently, some AI algorithms have taken steps to forget certain aged information in order to improve predictions.

[2026] This one held up. “Archival storage” legislation is mandatory data expiry.

  1. Privacy allows individuals to be authentic and to play appropriate roles in various contexts.

This basically states that how I behave and who I am depends on the context in which I am acting. There will be changes in this mutability since the context, often other people’s image of you, will be better informed of your overall role. Now when we enter a new context, we often assume other people have little to no knowledge of us and this allows us to develop our relationships unimpeded. This may or may not be true.

I think that we currently enter new contexts with uncertainty about what others know about us. With more readily available information, we could enter a new context with less anxiety, presuming that they already know some things about us but are willing to judge us based on our new context. In other words, I do not think expanding data access degrades this privacy.

  1. Privacy supports intimacy and the building of relationships.

There may be a small effect upon this reason. It will be easier to find information about a person but the information is not imposed into the relationship.

  1. Privacy supports the common good.
  2. Privacy protects from power imbalance between individuals and government/organizations.
  3. Privacy supports democracy, political activity, and service.

Quality data collection should improve these social ends. In fact the expansion of quality data is intended to improve these social ends.

[2026] Under universal read access, these three invert into a coercion machine. See the afterword.

  1. Privacy provides space in society for disagreement.

This is closely related to point number 4 and I think may be treated similarly.

Privacy is mutable. Making some changes in our perceived privacy will have benefits that outweigh the costs.

AI is the game changer. With data we gain the power to effect change and advantage in our world.

[2026] AI was indeed the game changer — in the opposite direction. See the afterword.


Afterword: The Construction Log (2026)
#

I did not just propose the lake. I tried to build it. This afterword records what the building taught me.

Every honest attempt at the design produced the same result. Health and financial data could not be openly readable, so they became credentialed sub-lakes. Credentials required an issuer, so oversight bodies appeared. Some records could not be safely exposed at the individual level at all, so aggregation layers appeared. With each iteration, the public data lake with equal access looked less like a commons and more like a tiered structure with sealed floors. I kept experiencing this as failure, as compromise of the vision. Even my own fixes did the same thing. To offset self-censorship, I reached for “strong anonymous channels”. An anonymous channel is privacy. I could not describe a livable version of my own proposal without rebuilding the thing I was abolishing. Eventually I stopped working on it, at exactly the paragraph above where the questions outnumber the answers.

There was a second, quieter, failure. The lake was supposed to produce quality data and it cannot. That was its entire justification. Erving Goffman drew the distinction via the concept of frontstage/backstage. People live a frontstage life, performed for an audience, and a backstage life where the performance drops. A lake everyone knows about converts all of life into a frontstage. What the data captures is the frontstage, performed life. This is corrupted data. Known observation corrupts the data at the source, at every scale the observation runs. The only honest measurements of unwatched behavior are those taken without people’s knowledge. That is data collection without consent, and the data-capture-free zone described below explicitly forbids it.

This corrects my responses to the first three of Magi’s fourteen reasons. Pretending they were not relevant actually harmed the data lake with corrupted-at-source data. My framework only considered data-use harms, so I missed the being-watched harms. A lake everyone knows exists puts everyone permanently on a frontstage. That is a harm even if no record is ever misused. There is more to say about frontstage/backstage, especially about where they have moved in the years since.

Why did my proposed symmetry keep breaking? Because equal access was a proxy for equalized power, and symmetric access between unequal parties does not equalize power. Equal access for unequals gives the powerful more data and better tools to read it. What the engineering kept forcing on me was an asymmetric structure. That structure was one lake with tiers of access. A permissioned pool has depth and rules on the surface. The depth is capture; surface rules fail by breach, by subpoena, and by changed purpose. That is the failure.

Instead, the lake must die. In its place, we build four separate structures to replace it.

Public statistics. Aggregate truths open to everyone. This preserves what my original article got right about the uncounted: the missing and murdered Indigenous women who did not exist in any national statistic, the Milwaukee crimes that vanished from the record while the department published flawed numbers. The uncounted are invisible, and this invisibility migrates power from the uncounted to the powerful. Recording this data removes that invisibility. It allows justice for the uncounted. The statistics must exist and must be public.

Accountable institutional access. Purpose-bound, audited access for researchers and regulators. This is the credit/debit scheme from my proposal, matured into credentials that log who asked what and why. This layer governs data already in institutional hands — medical, financial, and transactional — where consent-at-collection has never held under leverage. This also governs Magi’s reason 5, the sorting harm. A universal lake makes categorization cheaper for every insurer, employer, landlord, and lender at once. The harm of sorting is not that the data exists. The harm is that institutions make consequential decisions with that data without accountability. The remedy is restrictions on use: purpose limits, anti-discrimination enforcement, and algorithmic accountability.

A data-capture-free zone for persons. This is not a better-guarded vault. It is the zone promised in the foreword: capture prohibited by law and prevented by design, on the model of the secret ballot. My draft was correct that waivable privacy is already lost. So in this zone capture must be non-waivable. Anything waivable will be demanded by everyone with leverage over you. The only data that survives power is data that cannot be surrendered because it does not exist. The zone matters most where the people’s political and intellectual formation happens: meetings attended, causes funded, and books borrowed. A populace whose associations are readable can be coerced at every point of leverage. The courts saw this when Alabama demanded the NAACP’s membership lists. I claimed quality data would improve democracy. A fully readable populace is a fully enforceable one.

The inverse lake. One lake survives, aimed at privilege. Privilege has spent decades widening the information gap, not in a coordinated project but as the aggregate of moneyed actors each buying what their position invites. It buys data brokers to see us, and shell companies and legal complexity to hide themselves. The effect is structural and the inverse lake reverses that structure. Institutions must be readable to the people they govern: decisions, contracts, lobbying, enforcement patterns, budget lines, and ownership traces. Here, capture and access are the point. Data is legible when an ordinary citizen can extract what it means for them without hiring an expert. A nine-hundred-page filing is disclosed; it is not legible. Disclosure alone is not enough. We must have legibility: real-time reporting, beneficial-ownership tracing, and plain-language summaries. Madison said popular government requires popular information. The inverse lake makes that requirement operational. Its full defense belongs to a separate essay: where opacity is legitimate, how legibility avoids becoming a tool of power, and how it scales.

One sentence from my 2020 draft identified the right variable: “It’s not really a matter of what information is out there but how consolidated and cohesive it is.” I then drew the wrong remedy from it. The problem is not solved by consolidating everything in public. It is solved by governing the direction of legibility. Power must become legible to the people while the people’s formation must remain opaque to power.

In 2020, AI was my reason to want more open data. Since then, AI has multiplied every watcher’s harm-capacity: de-anonymization of “anonymized” records, inference of sensitive traits from innocuous data, and linkage of formally unlinked datasets. The case for the open lake weakens on its own terms. In opposition, the protective machinery matured. In 2020 it did not visibly exist. It exists now: privacy-preserving statistics of the kind the 2020 US Census deployed, proofs that verify claims about data without exposing the data, and computation over records nobody reads. The mathematics is real. Its full treatment lives in a later essay. The remaining questions are the ones that were always political rather than technical.

I proposed a symmetric data lake and then tried to build it. The engineering forced asymmetry on me because asymmetry is the proper design.


References
#

Added in 2026
#

  • Bernstein, Ethan S., and Stephen Turban. “The Impact of the ‘Open’ Workspace on Human Collaboration.” Philosophical Transactions of the Royal Society B, vol. 373, no. 1753, 2018, 20170239. https://doi.org/10.1098/rstb.2017.0239.
  • Goffman, Erving. The Presentation of Self in Everyday Life. Anchor Books, 1959.