# World Models and Proposition Search

The previous chapter assembled people, institutions, relationships, and world-state into a dyadic actor; this chapter now turns the construction into a recursively usable world model, which means the model must do more than predict one visible label, it must return a next state of the same type required by the following transition, otherwise the alleged world model is merely a decoder with delusions of grandeur.

The operational loop is repeated here because it is the object being constructed rather than merely summarized.[^recap-operational-loop]

<a id="gids-e3-recap-world-model"></a>
**GIDS–3 — Estimation, simulation, and trace, restated.**

\[
\mathcal H_{<t}\longmapsto\widehat D_t,
\qquad
(\widehat D_t,X_t=x_t,\boldsymbol\Xi_{t+1}=\boldsymbol\xi_{t+1})
\longmapsto
\widetilde D_{t+1}
\longmapsto
\widetilde O_{t+1}.
\]

The history available before the decision produces an estimated state; a candidate proposition encounters that state under a declared external scenario; the model simulates a next state and then the traces likely to escape. A hat means estimated from evidence, a tilde means simulated, and confusing the two is how a forecasting system begins hallucinating its own outputs into the database.

## Towards a Universal State-Transition Grammar

The ideal object remains the next phenomenal state. Let the complete actor-world state be

\[
\Sigma_{i,t}^{\star} :=
(T_{i,t},\phi_{i,t},c_{i,t},w_t),
\]

and recall the asymptotic silhouette,

\[
\phi_{i,t+1} =
F_i
\!\left(
T_{i,t},
\phi_{i,t},
 c_{i,t},
 w_t,
\mathbf p_{i,t}(x_t)
\right).
\]

If the complete state and exact law were available, the next phenomenal state would follow; the operational model possesses neither, therefore it works with an estimated state and a conditional distribution,

<a id="gids-e18"></a>
**GIDS–18 — Recursively usable state transition.**

\[
\widetilde s_{i,t+1}
\sim
K_{s,\theta}
\!\left(
\cdot
\mid
\widehat s_{i,t},
X_t=x_t,
\boldsymbol\Xi_{t+1}=\boldsymbol\xi_{t+1}
\right).
\]

The model receives the estimated actor-state, candidate proposition, and supplied or modeled exogenous change, then returns a distribution over the next state of the same operational type. This closure matters; a system predicting only “reply” or “no reply” cannot roll itself forward, because the next proposition will meet an actor changed by the first interaction, not the actor-state that existed before it.

The outward event is a partial consequence of the transition,

\[
\widetilde O_{t+1}
\sim
R_{O,\theta}(\cdot\mid\widetilde s_{i,t+1}),
\]

where reply, purchase, rejection, delay, concession, departure, promotion, thought expressed in language, or physical movement is the part that escaped into observation rather than the complete state itself. If the simulated next state does not screen off the previous state and proposition well enough, the trace law can condition on all three; that is an empirical choice, not a theological dispute about Markovity.

## The Proposition Must Be Rebuilt Inside the Actor

Begin with an external representation of the proposition,

\[
\mathbf e_t^x=E_{x,\theta}(x_t),
\]

which may encode language, structure, source, timing, price, channel, physical arrangement, or whatever features the system possesses; this remains the proposition as represented **outside** the receiving actor.

For dyadic state \(\widehat D_{ab,t}\), construct the receiving actor's proposition-relative representation,

\[
\mathbf p_{b,t}^{(D)}(x_t) =
\mathcal P_{D,\theta}
\!\left(
\mathbf e_t^x,
\widehat D_{ab,t}
\right),
\]

then form the interaction,

<a id="gids-e19"></a>
**GIDS–19 — Actor-relative proposition interaction.**

\[
\mathbf h_{ab,t}^{(D,\mathrm{int})}(x_t) =
\Psi_{D,\theta}
\!\left(
E_{D,\theta}(\widehat D_{ab,t}),
\mathbf p_{b,t}^{(D)}(x_t)
\right),
\]

and predict

\[
\widetilde D_{ab,t+1}
\sim
K_{D,\theta}^{\mathrm{int}}
\!\left(
\cdot
\mid
\mathbf h_{ab,t}^{(D,\mathrm{int})}(x_t),
\boldsymbol\xi_{t+1}
\right).
\]

Read the construction from left to right: encode what was presented; reconstruct it through the receiving actor and surrounding dyad; model the interaction between that reconstructed proposition and present state; simulate what the dyad becomes next. Role is not appended after the proposition has already been understood, because the role, relationship, institutions, and current state participate in the understanding itself.

This is the practical meaning of making realities composable. The common arena does not make sender, recipient, company, price, sentence, thought, and outward action interchangeable; it gives each a typed route into one transition grammar, where thought and movement remain two possible continuations of phenomenal state rather than two disconnected ontologies.

Two physically different propositions may be equivalent for one actor and task when they produce the same interaction representation. A percentage discount and corresponding dollar discount may be economically identical yet psychologically different; two messages may be linguistically different yet psychologically identical because the actor compresses both into “vendor asking for more of my time.” Equivalence belongs to the actor-relative transition, not the physical proposition alone.

## Why Sales Became the First Laboratory

The first organizational system attempted something broader and less measurable: ask a corpus of employees for opinions, loosely embed the people and decision-space, weight their judgments into swarm intelligence, implement one proposal, then wait for the organization to reveal whether the collective answer was wise. It failed because opinions carried weak consequences, justifications were cheap, decisions dissolved into the surrounding company, and the measurement process itself imposed operational labor; the full account appears in the opening because the failure explains the architecture that replaced it.[^failed-swarm]

Sales became the first laboratory not because selling exhausts the theory, nor because human decision is deterministic, but because it offers a harder training instrument than organizational opinion. A proposition can be timestamped, the actor and institutional context can be reconstructed from prior traces, a response and relationship update can be observed, and economic value can eventually be attached to the trajectory; the environment remains noisy and complex, which is exactly why the model carries actor, relationship, company, and world-state instead of asking participants to predict themselves.

## A Worked Dyad

Suppose a founder, \(a\), presents a partnership to an executive, \(b\). The executive's slow state contains durable estimates—tolerance for ambiguity, status posture, temporal preference, skepticism toward founder-led vendors, and other unnamed distinctions inferred from prior traces—while the fast state contains what has become active recently: a failed implementation, board pressure, budget deadline, irritation with the sender, perhaps a newly urgent internal problem. The company-state contains authority, incentives, institutional memory, current priorities, and who can veto the decision; the relationship-state contains prior promises, response rhythm, trust, and whether the founder has already exhausted the executive's patience.

Compare two propositions. The first asks for a broad strategic commitment, the second asks for one narrow technical test; an abstract classifier may see the same product and target, while inside the executive the first recruits loss of control, political risk, and implementation memory, and the second recruits curiosity, reversibility, and a chance to gather evidence without public commitment. The model is not merely choosing the better sentence; it is predicting two different state transitions.

After the actual response arrives, relationship and fast state change. A polite refusal may lower immediate probability while increasing trust; a meeting acceptance may look mechanically positive while revealing that the executive delegated the matter to someone without authority. The next proposition must meet the updated dyad, not the state that existed before the response.

## Simulation Is Not Filtering

<a id="gids-e20"></a>
**GIDS–20 — Simulation and filtering.**

Before the future is observed, the model simulates:

\[
\widetilde D_{ab,t+1}
\sim
K_{D,\theta}
\!\left(
\cdot
\mid
\widehat D_{ab,t},
X_t=x_t,
\boldsymbol\Xi_{t+1}=\boldsymbol\xi_{t+1}
\right).
\]

After real records arrive, the model filters:

\[
\widehat D_{ab,t+1} =
\mathcal F_\theta
\!\left(
\widehat D_{ab,t},
 x_t,
\mathcal H_{(t,t+1]}
\right).
\]

Simulation asks what might happen; filtering changes what we believe because something did happen. A model-generated future should never be written back as though it were evidence, and an observed response should not be treated as though it were the full state; the simulator produces distributions, the filter consumes timestamped records.

For proposition sequence \(\mathbf x_{t:t+H-1}\) and exogenous path \(\boldsymbol\xi_{t+1:t+H}\), recursive application produces a trajectory law over future states and traces. The environment does not freeze because recursion is convenient; markets move, people leave, companies change, later propositions depend on earlier responses, and anything caused by the proposition belongs inside the transition rather than being smuggled into an exogenous background held fixed across candidates.

Delayed outcomes require their own declared regime. A ninety-day event depends on continuation policy, future propositions, external change, and censoring; it is not an immediate emission from tomorrow's state merely because the database stores it on the same decision row.

## Learning From What Escapes

The state should support more than one visible consequence. Primary outcomes and auxiliary probes—intermediate actions, objections, delays, institutional changes, or other traces—may be modeled jointly when their dependence matters, or through separate heads when only marginal prediction is required; separate heads do not magically define a coherent joint future.

The training objective may combine the relevant predictive losses, probe losses, regularization, censoring, and time-to-event terms, although the generic gradient-descent equation has been omitted because every reader who reached this chapter already knows how parameters move.[^training-objective] The important distinction is temporal: parameters learn across cases, filtering changes belief about this case, fast actor and relationship states update when new records arrive, and slow states refresh only when durable evidence accumulates.

Probe heads survive only when they improve transfer, calibration, or stability of the state; a decorative taxonomy of motives is not a scientific result.

## From Forecasting to Proposition Search

This is where I stop pretending the purpose of the machinery is to admire prediction metrics. The system should compare admissible propositions by their expected effects on future state and downstream utility; otherwise why the hell are we building it.

Let \(\mathcal X_t^{\mathrm{cand}}\) be the candidate set available at decision \(t\), let \(\mathfrak e\) declare the continuation policy, future candidate-set process, exogenous-path law, and outcome convention, and let \(U_\tau\) be a measurable utility over predicted trajectories.

<a id="gids-e21"></a>
**GIDS–21 — Predictive proposition value.**

\[
V_{\theta,\mathfrak e}^{\mathrm{pred}}
\!\left(x\mid\widehat D_t\right) :=
\mathbb E_{\theta,\mathfrak e}
\!\left[
U_\tau
\!\left(
\widetilde D_{t+1:t+H},
\widetilde O_{t+1:t+H},
\widetilde{\mathbf Y}_t
\right)
\mid
\widehat D_t,
X_t=x
\right].
\]

When the maximum exists over the candidate set, choose

\[
x_t^\star
\in
\arg\max_{x\in\mathcal X_t^{\mathrm{cand}}}
V_{\theta,\mathfrak e}^{\mathrm{pred}}(x\mid\widehat D_t).
\]

The equation says: simulate each available proposition under the same declared future regime, score the resulting trajectories, and choose the best supported candidate. Utility belongs to the operator or system using the world model; it is not a claim about the objective the actor internally optimizes, because I do not need to infer a universal fitness function, expected-free-energy objective, or hidden rational program inside the person. I need to estimate how the actor is likely to decide over a long horizon under different degrees of information, then evaluate those futures under a constrained combination of interests appropriate to the application.

Sales is one laboratory, not the definition of the machinery. The same form can rank educational interventions, negotiation moves, organizational policies, recruiting sequences, product experiences, care plans, or any other propositions for which actor transition matters.

## Sequences, Not Isolated Tricks

The best immediate proposition may be a poor first move in a sequence. A message producing no meeting may reveal uncertainty, a reversible test may change trust, a difficult question may expose the true veto, and a temporary concession may alter relationship-state enough to make a later demand possible; therefore credit belongs to the trajectory.

For a policy \(\pi\), planning length \(H\), and declared regime \(\mathfrak e\), the value is the expected sequence of step utilities plus terminal value under recursively simulated states.[^sequence-value] The conceptual rule is simple: do not copy one eventual success backward and award it independently to every proposition that preceded it; that is not learning, it is numerology with a CRM.

The sensible progression is modest:

1. rank controlled proposition families one step ahead;
2. compare short predefined sequences;
3. choose the next proposition after each observed response;
4. attempt longer policy optimization only when the state transition and data collection process deserve trust.

## Prediction, Ranking, and Control

Forecasting estimates what tends to follow the proposition actually delivered; model-based ranking simulates alternatives and orders them under the fitted world model; interventional policy improvement claims that selecting a proposition causes a better outcome, and that final claim requires an experimental or otherwise defensible identification design.

The distinction does not need forty pages of self-flagellation. Observational success licenses forecasting and simulation inside the observed support; causal swagger begins only when the data collection regime earns it.

Until then, leave the causal swagger out of it.

## What I Would Actually Build First

The theory does not require loyalty to one architecture. I would begin with components I can debug when calibration goes sideways at two in the morning: strong tabular models for durable and institutional features, a small recurrent or state-space model for chronological history, explicit relationship-state, typed proposition encoders, a simple interaction block, and completely separate simulation and filtering paths.

The proprietary advantage is unlikely to come from choosing the most fashionable sequence block; it comes from the ontology, event-clock, actor construction, source-aware traces, and discipline of discovering distinctions that transfer. A capacity-matched monolithic model should still see the same data, because if it wins, it wins; the explicit construction earns complexity through transfer, data efficiency, calibration, controllable recursion, or interpretability useful enough to change decisions.

No actor model remains correct forever. People change, companies change, roles change, the same distinction becomes active under a new regime, and a coordinate that once transferred may stop carrying signal; when error, calibration, or support deteriorates, reopen ontology discovery rather than merely refitting weights. The next chapter defines the event discipline, baseline opposition, and four tests that decide whether the construction survives contact with the future.

---

[^recap-operational-loop]: This restates [GIDS–3 — Estimation, simulation, and visible trace](00_Opening.md#gids-e3); it is repeated because this chapter constructs the loop in full, and no new mathematical claim is introduced by the recap itself.
[^failed-swarm]: See [The First Machine Failed, Which Was Useful](00_Opening.md#the-first-machine-failed-which-was-useful). The failure is repeated only in compressed form here to explain why the first world model uses observed consequence, explicit institutional state, and a hard event-clock rather than self-reported organizational opinion.
[^training-objective]: One implementation may minimize \(\mathscr J(\theta)=\sum_{\ell}\lambda_\ell^Y\mathscr J_\ell^Y(\theta)+\sum_m\lambda_m^Z\mathscr J_m^Z(\theta)+\lambda_{\mathrm{reg}}\Omega(\theta)\), with nonnegative weights and head-specific masking, censoring, or survival likelihoods. This is an implementation interface, not a mathematical contribution.
[^sequence-value]: One explicit form is \(J_{\theta,\mathfrak e}(\pi\mid\widehat D_t)=\mathbb E_{\theta,\pi,\mathfrak e}[\sum_{k=0}^{H-1}\gamma^k u_\tau^{\mathrm{step}}(\widetilde D_{t+k+1},X_{t+k},\widetilde O_{t+k+1})+\gamma^H V_\tau^{\mathrm{term}}(\widetilde D_{t+H})\mid\widehat D_t]\). The regime must also define the future candidate-set and exogenous processes; otherwise the sequence value is not a well-defined comparison.
