Powered by Smartsupp

We Know His House, But Not His Life

September 2026

 

By: John F Groom

A White Paper on the Strange Shape of Human Data

Consider someone I will call “Old Friend.” I have known Old Friend for most of my life. We grew up together. I knew his family and his father, whom I will call Woody. This is not an attempt to identify a stranger from fragments found on the Internet. I already know who he is. I also know the broad outline of his professional life.

Old Friend graduated from a major Virginia university in the early 1980s. He served as a pilot in the United States Air Force. After leaving the Air Force, he joined a major American airline. He remained with that airline for essentially his entire civilian career, became a captain, and flew large commercial aircraft. By any ordinary standard, he had a long and successful professional career in one of the most highly trained, regulated, and documented occupations in the world. Recently, I wondered whether he had retired. It seemed like an extraordinarily easy question to answer. It wasn’t.

We Can Find the House

What makes the difficulty interesting is not simply that information about Old Friend is scarce. Quite the opposite. A surprising amount of information about him can be found. With relatively little effort, public and commercially aggregated databases can identify where a person has lived. Property databases can reveal details about the house: when it was purchased, its assessed value, lot size, square footage, number of bedrooms and bathrooms, previous transactions, tax history, and sometimes photographs and floor plans.

Other databases can expose ages, possible relatives, telephone numbers, previous addresses, property ownership, and other pieces of personal information. In Old Friend’s case, aviation databases also provide very specific information about private aircraft he has owned or built. Historical records can connect his name to a particular experimental airplane, registration number, serial number, and date of registration. The result is a strange asymmetry. We can learn a great deal about where Old Friend sleeps, but it is surprisingly difficult to determine what Old Friend spent his adult life doing.

The Question That Should Be Easy

The question was not particularly intrusive: What airline did Old Friend fly for, what aircraft did he fly, and has he retired? We already knew enough to narrow the search dramatically. We knew his full identity, where he lived, approximately when he was born, and where and approximately when he went to college. We knew that he served as an Air Force pilot, subsequently spent decades flying for one major airline, eventually became a captain, and that aviation was not merely an occupation for him but a serious personal interest.

That ought to be sufficient. Yet ordinary Internet searches produced remarkably little. There were scattered references to an airline pilot of the correct name and location, aircraft-registration records, references connecting him to aviation organizations and to his father, and old newsletters and other fragments.

But the simple professional chronology remained elusive:

Air Force → airline → aircraft assignments → captain → retirement.

The information almost certainly exists. It just isn’t connected.

The Wrong Things Are Structured

This exposes something important about the modern data environment. The Internet does not necessarily contain the information that is most important about a person. It contains the information that institutions happened to collect, digitize, structure, publish, and make searchable.

Real estate records are highly structured because governments need them for taxation and ownership. Aircraft registrations are structured because aircraft are regulated. Addresses are aggregated because there is a commercial market for them. Consumer behavior is recorded because advertisers value it. Credit information is organized because lenders need it.

And so the databases become extraordinarily good at answering questions such as: Where does this person live? How much is the house worth? When was it purchased? What aircraft is registered in his name? Yet they may be terrible at answering: What did this person actually accomplish during forty years of professional life? That is a profound mismatch.

A Successful Career Becomes Data Exhaust

Consider what a career like Old Friend’s actually contains. Becoming a military pilot requires years of selection, education, training, and evaluation. Flying for a major airline requires additional training, certification, recurrent testing, and thousands of hours of experience. Becoming a captain requires still more experience and responsibility. Flying increasingly sophisticated aircraft over several decades represents an enormous accumulation of human capability.

There would have been training records, flight records, licenses, certifications, military assignments, airline seniority records, aircraft qualifications, recurrent evaluations, crew assignments, union records, company records, and eventually retirement records. In other words, Old Friend’s career may have generated millions of individual data points. Yet from outside those institutional silos, much of the career effectively disappears. We can see the house. We cannot see the life.

Privacy Makes the Irony Greater

There is another dimension to this problem. Some of the information that is readily available probably should not be so readily available. A person’s exact home address, detailed characteristics of his house, family associations, and other personally identifying information have obvious privacy implications. There are legitimate reasons to question why strangers should be able to assemble such information with a few searches.

Meanwhile, broad professional information is far less sensitive. Knowing that someone graduated from a particular university, served as a military pilot, joined a major airline, became a captain, flew particular categories of aircraft, and retired after a successful career tells us something meaningful about the person without telling us where his bedroom is. Yet our information infrastructure frequently makes the first category easier to obtain than the second. This is almost exactly backwards.

Search Is Not Knowledge

The Old Friend example also illustrates a misconception created by the Internet. Because we can search billions of documents almost instantly, it is tempting to believe that we have constructed something approaching a comprehensive database of human knowledge.

We have not. We have constructed an extraordinarily large collection of partially structured fragments. Search engines are exceptionally good at locating fragments when the right fragment happens to have been published and indexed. Artificial intelligence is increasingly good at interpreting and recombining those fragments. Neither solves the underlying problem when the necessary pieces remain isolated.

Suppose one database knows that Old Friend graduated from college in 1983. Another knows that a person with his name was an airline pilot. Another records his private aircraft. Another contains an old airline seniority list. Another contains his military history. Another contains a retirement notice. A human who already knows Old Friend can immediately understand that these facts probably describe a coherent life. The databases cannot necessarily do so. The problem isn’t merely retrieval. It is coordination.

Data Without Provenance Creates Another Problem

Simply connecting everything would create obvious dangers. There are many people with identical or similar names. In researching Old Friend, for example, another aviation professional with a similar name and a substantial Air Force career repeatedly appeared in search results. An AI system could easily merge the careers of two different people and confidently manufacture a biography that never happened.

So the solution cannot simply be: Collect more data and connect everything. A useful human-data architecture must preserve provenance. Every assertion should retain its source. Identity matches should have confidence levels. Contradictory information should remain visible. User-supplied knowledge should be distinguishable from independently verified information. Sensitive information should be treated differently from ordinary biographical information.

The system needs to be capable of distinguishing between what we know, what we have evidence suggesting, what someone who personally knows the individual reports, and what we do not yet know. That distinction becomes increasingly important as AI becomes capable of constructing extremely convincing narratives from incomplete evidence.

The Human Record Should Be Organized Around the Human

Most databases are organized around the needs of institutions. The tax authority organizes information around property. The airline organizes information around operations. The military organizes information around personnel and missions. The university organizes information around students and alumni. The FAA organizes information around aircraft and certifications. The bank organizes information around accounts.

Each system may work perfectly well for its own purpose. But the human being exists across all of them. From the individual’s perspective, these aren’t separate worlds. They are chapters of one life. Old Friend did not experience himself as a record in an Air Force database followed by an entry in an airline personnel system followed by an aircraft-registration record followed by a property-tax record. He experienced one continuous life. Our data architecture does not.

From Data Collection to Human Understanding

This distinction will become much more important in the age of AI. The first phase of the digital revolution was largely about digitization. The second was about search. The third has been about AI’s ability to understand and generate information. But a further step is necessary: turning fragmented data into coherent, provenance-preserving representations of reality.

For human beings, that means something closer to a longitudinal record. Not a giant surveillance file or an indiscriminate collection of everything that can be discovered about someone. Rather, a structured representation of the important events, capabilities, relationships, decisions, achievements, and changes that constitute a life, with appropriate consent, provenance, privacy, and access controls. The distinction matters. A good system might know that Old Friend spent decades safely flying sophisticated aircraft. It does not necessarily need to tell strangers the dimensions of his house.

What Do We Actually Want to Know?

Old Friend provides a useful test because the original question was so mundane. I wasn’t conducting an investigation, trying to locate him, or trying to discover something embarrassing. I simply wondered: Has my old friend retired? Behind that question was something more human: What happened to someone I grew up with? How did his career turn out? Where is he in his life now? Our information systems could tell me an astonishing amount about his property. They struggled to tell me the broad outline of the life I already knew he had lived. That is not primarily a failure of search. It is a failure of how we have structured information.

The Larger Lesson

We frequently describe the modern world as suffering from information overload. That is true in one sense. But in another sense, we suffer from something quite different: an abundance of data combined with a scarcity of organized meaning.

There may be thousands or millions of records associated with a human being. But those records were created for different purposes, stored in different systems, governed by different standards, and rarely designed to interoperate.

AI makes this contradiction increasingly visible because we can now ask natural questions that cut across those boundaries. What did this person spend his life doing? The question makes perfect sense to a human, but it may make almost no sense to the databases.

The next generation of information systems should therefore not merely ask how much data we can collect. They should ask what we are actually trying to understand, and perhaps an equally important question: Why do we know so much about the things that should be private, and so little about the things that actually tell us who someone is?

Old Friend’s house is a data object. His career is a life. We have become remarkably good at documenting the first. The much more important challenge is learning how to represent the second.

 

Whether you're exploring interoperability, dataset valuation, AI readiness, or ecosystem participation, we welcome conversations with researchers, organizations, and strategic partners interested in the future of structured data systems.

info@datauniversa.com