Evaluation and Ontology Growth conceptual illustration
[↗] GIDS series

preprint · chapter 7 · 7 of 8

Evaluation and Ontology Growth

The philosophy explains why the model has this shape; data decides whether the shape survives. This chapter does not prescribe the company’s current benchmark, expose the discovered ontology, or freeze the research program around one early sales dataset, because the canonical manuscript should define the evaluation discipline while the empirical paper supplies one public implementation, one corpus, one set of results, and all the humiliating detail required for replication.

The operational loop has already been derived, therefore it is not re-derived here; the evaluation problem begins at the instant an estimated state is frozen before a proposition, then asks whether the predicted transition survives future time, strong baselines, transfer, and intervention.1

Event Time Is Sacred

The clock has to be clean or the entire evaluation becomes leakage with equations around it. Let index decision epochs and let be timestamped records; only records fully available before proposition was selected may enter the state used to forecast its consequences.

GIDS–22 — Decision-aligned event time.

The equation is small and carries the entire leakage argument. If an email response arrives at 14:02, it cannot appear in the state constructed for a message selected at 14:00; if a transcript summary is generated after a meeting, it cannot enter the pre-meeting actor vector; if a company fact is learned three weeks later, the historical model does not get to know it early merely because the warehouse does now.

At every decision epoch: construct the actor or composite state from pre-decision history; record the candidate propositions actually available; choose and deliver ; freeze the forecast; append observations only when they occur; attach delayed outcomes only when their horizons mature. Timestamp ties require an explicit order, because “same day” is not an event clock.

Decisions, Observations, and Outcomes

A decision record should preserve the decision identity, candidate set, chosen proposition, elapsed time, and—when policy evaluation is intended—the assignment probability or density and information actually available to the historical policy. Observations arrive later and may include outward response, new actor evidence, institutional change, or world-state change; each enters the filtered state only when it becomes available.

A decision may carry several outcomes at different horizons, and every outcome requires an availability or censoring indicator. An outcome that has not matured is not a negative; it is unavailable. Long-horizon evaluation should therefore use complete follow-up windows or appropriate time-to-event methods instead of converting ignorance into failure.

The manuscript remains vague about the content of discovered actor axes and application-specific fields; it should not be vague about time, masks, source, candidate availability, or outcome maturity. Secrecy around the ontology is defensible; sloppiness around the dataset is not.

Evaluate the Future, Not a Shuffled Past

Random row splits are nearly useless for repeated actor interactions, because they leak future actor-state into training, place the same relationship on both sides of the split, and reward memorization as though it were generalization. Use nonoverlapping temporal windows,

and make model choice, ontology changes, thresholds, and calibration decisions using training and validation only; the final test window remains untouched until the analysis is frozen.

Time is necessary and insufficient. Report transfer separately for later interactions with known actors, new actors inside known institutions, known actors under new roles or proposition families, entirely unseen actors and institutions, and new tasks or outcome horizons. The final regime matters most to the ontology claim, because a coordinate discovered for one target earns its place only when it preserves useful structure elsewhere.

Rows remain dependent within people, relationships, institutions, campaigns, and time periods; uncertainty estimates must respect those clusters. A million events generated by ten companies are not a million independent companies.

The Models That Must Be Beaten

Every implementation should face at least three levels of opposition. The first is a shallow current-proposition model using the obvious predictors available now, with no durable actor-state; the second is a strong structured-history model using static actor and institution inputs plus competent hand-built summaries; the third is a capacity-matched monolithic sequence model seeing the same pre-decision records without being forced to separate slow actor, fast state, relationship, institution, and proposition interaction.

The monolithic model is the dangerous one. If it consistently predicts better, transfers as well, calibrates as well, and requires no more data, the explicit decomposition has become a story told after the fact; if the structured model ties on raw prediction while transferring better, remaining stable under sparse observation, or supporting reliable recursion, that may still justify it, although the trade must be stated rather than smuggled in as “interpretability.”

Four Essential Tests

The earlier manuscript carried a ceremonial army of ablations; four are enough to decide whether the central claims are alive.

1. Slow versus fast state

Remove , then remove . Slow state should matter most under sparse observation, cold start, role transfer, and longer horizons; fast state should matter most after recent events and at shorter horizons. If their removal produces no distinguishable pattern, the decomposition is not buying what it claims.

2. Actor, relationship, and institution

Replace the structured composite state with flat metadata and shallow history; this tests whether relationship and institutional access structure contain signal that cannot be reduced to ordinary tabular features.

3. Explicit construction versus monolithic sequence

Give both models the same information and comparable capacity. This is the shortest route to discovering whether the ontology is useful or merely narratively satisfying.

4. Cross-task transfer

Discover or fit actor structure on one family of outcomes, then evaluate whether it improves another without being rebuilt from scratch. This is the decisive test for a new axis; a coordinate that helps only the target that created it is probably a target feature wearing philosophical makeup.

A fifth diagnostic is often cheap: shuffle recent within-actor or within-relationship history while preserving static profiles. If performance does not fall, the model was not using sequence in the way claimed.

How a New Distinction Earns Its Place

Ontology growth begins with structured residual error. Suppose the current registry systematically fails for a subset of actors, roles, propositions, or regimes; propose an operation —split, merge, rotation, interaction, source separation, role restriction, new family, weakening, or retirement—then ask whether it explains a stable residual pattern in training, improves future risk in validation, and transfers across held-out tasks, roles, horizons, or populations.

The transfer score is repeated from the ontology chapter because it becomes the evaluation rule here.2

GIDS–7 — Cross-task ontology retention, restated.

Retain the distinction when the score is positive, the gain repeats across future time, and—where intervention data exists—the distinction participates in the predicted change under a controlled proposition. The name may change later; the predictive relation is what survived.

This is the model-side analogue of evolutionary accumulation, not the same mechanism. Evolution creates possible distinctions in organisms; ontology growth creates typed hypotheses about distinctions in the model, and the hybrid registry remains useful because it can be debugged even if a continuously updating manifold eventually captures the underlying geometry more faithfully.

Policy Evaluation Without Causal Swagger

When assignment probabilities were logged at decision time and outcomes have matured, off-policy estimators may reweight historical decisions toward a target policy; the standard one-step inverse-propensity expression is retained in a footnote because it is an evaluation instrument, not part of the philosophical core.3

The conditions matter more than the formula: overlap, correct event ordering, stable treatment definition, defensible assignment, mature outcomes, censoring discipline, and support. Free-form propositions generally possess almost no exact historical overlap, therefore early policy evaluation should operate over controlled proposition families or dimensions instead of pretending every new paragraph has a counterfactual twin in the logs; sequential policies require sequential estimators, and one-step arithmetic should not be stretched across an entire conversation.

Metrics Should Match the Claim

Use proper scoring rules for probabilistic forecasts, calibration measures for decision use, ranking metrics for candidate ordering, and survival or competing-risk methods for censored outcomes. No single number is the research program.

GIDS–23 — Future-risk criterion.

Choose before the final test window. “Statistically detectable” is not automatically “worth the machinery”; report gain size, uncertainty, calibration, data requirements, and transfer, while predictive ranking should also be checked for support because an optimizer will discover propositions the model likes for accidental reasons.

Drift Is Not an Exception

Recent risk, calibration, support, source mix, and event quality should be compared against reference windows. Persistent degradation reopens the model: refit when parameters moved, reopen ontology discovery when residual structure changed, revise the actor boundary when the old scale no longer explains the trace, and update institutional access when information begins flowing through another authority structure.

The model is allowed to become wrong; it is not allowed to become wrong silently.

What Counts as Success

Success is not that the paper sounds deep. The program succeeds operationally when an explicit actor-state model improves future prediction or transfer over strong baselines; when slow and fast components fail in the different regimes their meanings predict; when actor, relationship, and institutional construction add reusable information; and when newly discovered distinctions generalize beyond the labels that proposed them.

It succeeds more strongly when the same state supports several decisions over long horizons and under different degrees of available information. The program-level stopping rule remains deliberately unromantic: if three materially different implementations of the explicit decomposition, evaluated across at least two independent temporal corpora, fail to match a capacity-comparable monolithic sequence model and produce no repeatable cross-task transfer, retire the decomposition for that actor class; keep the useful event discipline, stop calling the latent construction a stable ontology.

That rule is not the center of the manuscript, it exists so failure has somewhere to land. The research program remains open-ended because reality keeps producing distinctions, not because every failed implementation receives an infinite appeal.


In Memory of Einar Kringlen

Einar Kringlen seated in his office in front of bookshelves.

It has been an honor tackling this multi-generational problem with you; to you I owe much.


Footnotes

  1. The operational loop is GIDS–3, constructed in full in World Models and Proposition Search; this chapter assumes that definition and evaluates its consequences rather than deriving it again.

  2. This is GIDS–7 — Cross-task ontology retention, repeated because the evaluation chapter must state the acceptance rule locally; no new mathematical claim is introduced.

  3. For eligible decisions , logged assignment probability or density , mature utility , and target policy , one basic estimator is . It is valid only under the stated overlap, logging, ordering, treatment, and censoring conditions.