Skip to content

Egeria Resource Explorer (Part 2): The Anatomy of a Question

Active Surveying Beyond the Catalog Perimeter · Part 2

One catalog of questions across five resource types, and what a single row of it actually contains.

If the Funnel Is the Shape, This Is the Unit

In Part 1 we argued that exploring things outside the catalog perimeter means keeping surveying separate from cataloging, and laid that out as a funnel with six stages: Scouting, Discovery, Assessment, Analysis, Enrichment and Curate. This post is about the thing the funnel is made of.

A six-stage funnel narrowing candidate code repositories: Scouting, Discovery, Assessment, Analysis, Enrichment, and Curate and graduate. Each stage exchanges data two-way with Egeria. Three capabilities wrap around the funnel without being stages: Investigation is the scope the run lives in, Understanding is a lens onto everything measured so far, and Automate drives scheduled re-runs. Graduated assets pass through an outbox into the Egeria enterprise catalog and stay under continuous surveillance.
The survey funnel, from Part 1. Six stages narrow the candidates; Investigation, Understanding and Automate sit around them rather than in them.

A note on these blogs. We’re building Resource Explorer as we write this series, so the measurements and limits described here were true when we wrote them, and we expect many of them to be out of date quickly. That’s deliberate. We want to be honest about the gaps we find, because finding an issue is the first step toward fixing it. So read this as a view from the inside — what we think is good, as well as where we know it falls short. Progress marches on.

Think about what that funnel is for. Say a department is running eighty databases nobody has written down, a few hundred internal Git repositories, and filesystems that have been growing for a decade. None of it is in a catalog, and a good deal of it doesn’t deserve to be. You can’t examine all of it properly. Profiling every schema, classifying every column and parsing the source of every repository simply isn’t something you can afford to do. A cheap first sweep separates the dead staging instances from the dozen candidates worth a closer look.

The trouble is that governance tooling has a habit of turning a useful heuristic into an assembly line. Six stages become six gates, and someone who already knows which resource they care about still has to click through all of them to get to it.

There’s a distinction worth drawing here, because it decides what this tool is for. Two jobs often get confused with each other: surveying resources that aren’t in the catalog yet, and performing analyses on those that already are. Resource Explorer is built for the first — the frontier. We may extend it to do more of the second, and some of what it records is useful there, but scouting the unknown is the job it’s shaped around.

Even at the frontier, people don’t ask questions in order. Someone evaluating an unfamiliar library wants to know straight away whether it carries unpatched security advisories, before anything else. Someone handed a database they’ve never seen wants to know what’s actually in it — what the tables mean, where the data came from, what it’s safe to use for — and not only who owns it. Someone looking at a shared filesystem wants to know whether sensitive files are sitting somewhere they shouldn’t be. Every one of those is a question from the middle of the funnel, asked first. The catalog records which stage each question belongs to and the interface shows it, so asking out of order is a choice you can see yourself making rather than an accident.

So we treat the funnel as a suggestion rather than a rule. The unit that matters is the question. People arrive with questions, and our job is to give them useful answers; everything else here is machinery in service of that. The funnel proposes a sensible order for narrowing a field of candidates, but the tool behaves more like a workbench, where you can ask a particular question whenever you want to ask it.

What a Question Is Made Of

A stage doesn’t run a fixed list of checks. It asks questions, and each question carries everything needed to answer it. Here’s one, verbatim from the catalog:

Who owns this resource (accountable owner), and who administers it?

One row of the question catalog, split in two. Five parts of the row — the question text, why it matters, its rationale, its funnel stage and its perspectives — are published into Egeria as a Question glossary term. The rest — resource types, level, answering analysis, answering mechanism and purposes — have not crossed into Egeria yet. The answering analysis shown, repository_health plus chaoss_metrics, answers the question for repositories only; other resource types need different analyses. The boundary is a snapshot of a migration, not a fixed design.

Who owns this resource, and who administers it? We ask it at Scouting, because cheap signals can answer it. It applies to all five resource types, because ownership is the same question whether you’re asking it of a repository or a database. But the same question doesn’t mean the same answer. For a repository, two analyses answer it: repository_health, which tells you whether the thing is maintained at all, and chaoss_metrics, which tells you who maintains it and how few of them there are. A database needs a different answer from a different source. The question is shared; the machinery underneath it isn’t. The row also records two of the ten purposes it serves, Explore and Select, and three of the twelve perspectives it answers to: Steward, Data Owner and Security.

It’s worth being honest about what that answer is worth. For a repository you get the organization behind the GitHub account, which may be a name nobody uses any more — Egeria’s own code sits under an odpi org for an organization that no longer exists. For a database you get the role that owns the objects: a user id, with nothing about the person behind it, the team they sit in, or what they’re accountable for. The question is answerable; the answer is thinner than it looks.

The rest of it usually exists somewhere — a project website, internal documentation, or someone down the hall who simply knows. Often you have to look under a few rocks to find it. That is why Enrichment carries an Owner field of its own, where a person records who actually answers for the resource. The two ends are not yet joined: the contributor data gathered at Scouting does not flow into that field, and the interface says so directly underneath it — offering instead to let someone stand as interim owner until the real answer turns up.

The catalog also has room to record its own shortfalls, in the same row as the question rather than in a backlog somewhere else. One row still reads Status — inadequately answered today, against the question of what languages and file types make up a repository. The question is asked, an analysis answers it, and the catalog says plainly that the answer isn’t good enough yet. A shortfall recorded in the open is also an invitation. If you care more about a particular question than we do, you can see exactly where our answer falls short. What you’d extend isn’t really the catalog but the analysis behind it: the question stays as it is, and the work goes into tuning or replacing the analytics inside Resource Explorer that answer it, until the answer meets your requirements rather than ours. That’s one of the ordinary advantages of doing this in the open.

The rows below are the catalog as it stood the day we wrote this, not a finished state. The similarity gap, for one, closes as soon as we load the vector store.

QuestionStageTypesAnswering analysisPerspectives
Who owns this resource (accountable owner), and who administers it?Scoutingallrepository_health + chaoss_metricsSteward, Data Owner, Security
What are similar resources, and how does this differ?AnalysisallGAP: similarity search over pgvector embeddings (not yet built)Community, Security, Admin
What does this resource cost to run, host or license?Analysis / EnrichmentallHuman-supplied via the Enrichment Context formFinancial, Admin
Is this database alive — writes since the statistics were reset, last vacuum or analyze, and is anything reading it?Scoutingdatabasedb_activity_signalsSteward, Data Owner, Admin
Four rows from resource_questions.csv, abridged to five of its twenty-two columns.

One Catalog, Five Resource Types

The questions live in a single spreadsheet — one CSV with a Resource Types column, rather than one file per technology. The reasoning is simple. “Who owns this data?” is the same question whichever resource you’re asking, and authoring it five times is five chances to drift. Values are repo, database, filesystem, dataset, model, or an asterisk for all of them. If that column contains something the generator doesn’t recognize, it stops with an error rather than quietly producing nothing — because a misspelling like databse and a resource type nobody has written questions for yet would otherwise look exactly the same.

Each question is then organized along four further dimensions:

  • Funnel Stage (the suggested order) — Scouting, Discovery, Assessment, Analysis or Enrichment, balancing computational cost against analytical depth. A stage may be slash-combined (Analysis/Enrichment) to signal an analysis-first, human-fallback pattern.
  • Level (how deep to look) — resource, container, member or field. The vocabulary is deliberately independent of technology: a database’s container is a schema, a repository’s is a module. That way “who is allowed to read this?” can be one question asked of databases, filesystems and datasets alike, rather than three questions that have to be kept in step. A question can carry more than one level, and is then offered at each of them. “How big is this database?” is both resource and container, because the answer is a total and a breakdown at once.
  • Twelve Perspectives (viewpoints) — Financial, Governance, Steward, Data Owner, Consumer, App/AI Builder, Privacy, Community, Data Expert, Security, Architecture, Admin. These are lenses, not job titles: one reviewer holds several at once and switches between them depending on what they’re doing that morning. They also do something more practical. When a resource has dozens of questions against it, the whole list is overwhelming; choosing a perspective is how you narrow it to what you care about right now.
  • Ten Purposes (what you’re trying to do) — Explore, Select, Assess, Certify, Deploy, Maintain, Learn, Share, Remediate, Attest. A team evaluating a library to Select asks early qualification questions; an engineering team trying to Maintain an existing database asks about schema stability and operational risk. Purpose narrows the list the way perspective does, and it also shapes what a useful answer looks like: the same finding is read differently by someone deciding whether to adopt a thing and someone who already depends on it.

Level is what lets the funnel narrow twice. The stages narrow a population of candidates; the levels narrow the scope inside whichever candidate survives — the whole database, then a schema, then a handful of tables. A four-hundred-table warehouse that answers every question as a single rollup isn’t really being surveyed, it’s being summarized. We’ve authored the levels. Most of the analyses behind them still run whole-resource, which is the next piece of work rather than something we can claim is done.

The later stages are what give that second narrowing its weight. Enrichment indexes content so you can find it later by meaning rather than by name — the step behind retrieval-augmented generation — and it isn’t free. It costs processing on the way in and storage from then on. How much there is to index depends on the resource. A repository is mostly text, so a fair amount of it can reasonably go in: source, documentation, the READMEs people actually wrote. A database isn’t text, and what’s worth indexing is much narrower — the documentation and the metadata around the tables, not the rows inside them. We’re not going to load an entire database into a vector store. This is the stage where the bill scales with how much you chose to keep, which is the argument for narrowing to a few schemas before you get here.

That coverage is broad but uneven. Many questions apply to every resource type, which is why datasets and models have real numbers against them rather than zeros. But asking the same question everywhere doesn’t mean answering it everywhere. “Who owns this?” is a single row whichever resource is asking, and answering it may still take a different surveyor for each type, or one surveyor generalized until it reaches them all. The questions span five types today. The analyses behind them run out well before that.

One stage is missing from that list, and its absence is deliberate. Curate doesn’t ask the resource anything, so it has no questions of its own. We come back to it in Part 3.

Looking Ahead to Part 3

So far this is a question at rest: a row in a spreadsheet with everything needed to answer it written down beside it. Part 3 is about what happens when someone asks one. A question near the bottom of the funnel usually depends on work nobody has done yet, and that raises the two problems worth the rest of this series: how the system works out what has to run first, and how it knows what that will cost before it spends your afternoon on it.

Links

As always, feedback is welcome — and if you’re interested in adding your favorite questions (and maybe some new analyses), we’d be glad to welcome you to the project.