# Predictive Actor-State

The Chimera chapter separated the slowly realized actor from fast local state, then made role and proposition part of the interaction rather than treating a person as one timeless factor score; this chapter now asks the more severe question, which is not whether the complete mental interior can be recovered, but what information must survive compression if we want to predict the actor under a declared family of future propositions.

The operational state introduced earlier is repeated here because the predictive construction must remain legible without sending the reader backward through another file.[^recap-operational-state]

<a id="gids-e11-recap-predictive"></a>
**GIDS–11 — Slow and fast operational state, restated.**

\[
\widehat s_{i,t} =
\left(
\widehat{\mathbf t}_{i,t},
\mathbf z_{i,t},
\mathbf c_{i,t},
\mathbf w_t
\right).
\]

This is the finite object estimated from traces; it is not the complete phenomenal state, nor is it automatically sufficient merely because we gave its components impressive names.

## The Notion of State

Philosophically, everything belongs inside state. The actor's present state includes perception, interoception, memory, attention, anticipation, language already moving through the mind, action tendencies, action already underway, and whatever else is available to the actor at that instant; call the complete phenomenal state

\[
\phi_{i,t}\in\Phi_i.
\]

That is the motivating object, and it is inaccessible. A transcript is not the thought that produced it, a click is not the desire, a psychometric score is not the person, and even perfect outward logging would leave many internal routes observationally equivalent; therefore, if I point the engineering section directly at \(\phi_{i,t}\), I begin lying almost immediately.

The operational question is instead:

> What must be preserved about the actor's available history so that future response under admissible propositions can be predicted?

This question has a mathematical answer even when the complete interior does not.

## The General Predictive Response Object

Define the ideal pre-proposition information state

\[
\mathsf I_{i,t} :=
\left(
\mathcal H_{i,<t},
T_{i,t},
 c_{i,t},
 w_t
\right).
\]

The terms are the history available before decision \(t\), the slowly changing realized actor, present context, and the relevant world; the complete phenomenal state is not granted to the predictor, because if it were, much of the problem would disappear by definition.

For a finite horizon \(H\), let \(\mathbf x_{t:t+H-1}\) be an admissible sequence of propositions and \(O_{i,t+1:t+H}\) the future observable traces. The response law is

<a id="gids-e12"></a>
**GIDS–12 — Proposition-conditioned response law.**

\[
\mathscr R_{i,t}^{(H)}
\!\left(
\cdot
\mid
\mathsf I_{i,t},
\mathbf x_{t:t+H-1},
\boldsymbol\xi_{t+1:t+H}
\right) :=
\mathcal L
\!\left(
O_{i,t+1:t+H}
\mid
\mathsf I_{i,t},
\mathbf X_{t:t+H-1}=\mathbf x_{t:t+H-1},
\boldsymbol\Xi_{t+1:t+H}=\boldsymbol\xi_{t+1:t+H}
\right).
\]

The notation is dense; the idea is not. Hold one actor-information state fixed, present a possible proposition sequence, supply or model a future external scenario, and ask for the distribution of traces that follow. The admissible family must be declared, because “every possible proposition under every possible universe” is not a usable scientific object.

Two information states are predictively equivalent when every admissible proposition sequence produces the same family of response laws:

<a id="gids-e13"></a>
**GIDS–13 — Predictive equivalence.**

\[
\mathsf I\sim\mathsf I'
\iff
\mathscr R^{(H)}(\cdot\mid\mathsf I,\mathbf x,\boldsymbol\xi) =
\mathscr R^{(H)}(\cdot\mid\mathsf I',\mathbf x,\boldsymbol\xi)
\]

for every declared horizon, proposition path, scenario regime, and measurable future event in the family under study. The equivalence class is the ideal **general predictive actor-state**, denoted \(Q_{i,t}\).

This is the cleanest formal answer to the state question. If two actor histories differ in a thousand details but imply the same response law under every proposition we care to present, those differences do not belong in the minimal predictive state; if one forgotten humiliation changes one response family ten steps later, it belongs. The object may be infinite-dimensional, and there is no promise that every actor can be compressed into one convenient finite vector without loss; the paper becomes much more honest the second this object, the phenomenal state, and the estimate stop being treated as the same thing.

For a narrower task \(\tau\) and horizon \(\Delta\), a task-conditioned summary may exist:

\[
q_{i,t}^{(\tau,\Delta)} =
\Pi_{\tau,\Delta}(Q_{i,t}).
\]

The map keeps what one task and horizon require, and need not be linear; a seven-day response prediction may discard structure required for a five-year succession decision. For every measurable outcome event \(B\), the summary is sufficient when

<a id="gids-e13a"></a>
**GIDS–13A — Task-conditioned sufficiency.**

\[
\mathbb P
\!\left(
Y_{i,t}^{(\tau,\Delta)}\in B
\mid
\mathsf I_{i,t},X_t=x
\right) =
\mathbb P
\!\left(
Y_{i,t}^{(\tau,\Delta)}\in B
\mid
q_{i,t}^{(\tau,\Delta)},X_t=x
\right).
\]

Once the proposition and task-summary are known, the richer information state contributes nothing further to this outcome under the declared regime; the causal version replaces observational conditioning with the corresponding intervention-indexed laws. The general state \(Q_{i,t}\), however, is judged against the broader declared family of admissible proposition-conditioned response laws, which is where cross-task transfer enters at the foundation rather than as a decorative secondary metric.

## Predictive State Representations, and the Difference

This construction is closely related to predictive state representations: represent a latent dynamical state through predictions of future observable tests under possible actions.[^psr] The resemblance should be stated because it gives the manuscript a real mathematical neighbor; the difference is not that PSRs manipulate vulgar worldly objects while GIDS manipulates holy mental ones, the difference lies in the actor-relative construction placed before predictive state.

In GIDS, the externally modeled proposition, the proposition reconstructed by the actor, the conscious or actor-available state, the memory it recruits, the thought or movement it produces, and the state that follows all enter one actor-relative transition grammar. We keep types so the implementation does not perform nonsense arithmetic, but we do not place an ontological abyss between observation, thought, feeling, and action once they are inside the actor's phenomenal loop; each is an organized difference changing what becomes available next, with outward action distinguished mainly because another observer can register more of its consequences.

The second difference is transfer. A narrow predictive state may be sufficient for one controlled dynamical system; GIDS wants a reusable actor ontology, a representation of response structure surviving changes in proposition family, role, horizon, and task. That ambition is harder and may fail, although it is also the reason the latent object matters more than one forecast. No theorem is inherited merely because the objects look related; existing results apply only when their assumptions match the actor system constructed here.

## Minimality Without a Soul Coordinate

Call a general predictive state \(Q_{i,t}\) minimal when every other state \(R_{i,t}\) preserving the same declared family of proposition-conditioned response laws contains enough information to recover it:

\[
Q_{i,t}=h(R_{i,t})
\quad\text{almost surely}
\]

for some measurable map \(h\). This does not identify one sacred coordinate chart; rotations, invertible transformations, and more complicated reparameterizations may preserve every relevant response law, therefore the ontology is not the literal name attached to each axis, it is the stable structure of distinctions and relations needed to preserve response across tasks.

Human-readable names remain useful handles. They are not divine certificates.

## What Approximation Means

The model estimates \(\widehat s_{i,t}\) from records available before the proposition, while the ideal information state remains richer. Approximation should therefore mean more than drawing a wavy line between two symbols.

For outcome \(Y_{i,t}^{(\tau,\Delta)}\), define the information lost by state compression as

\[
\epsilon_{\tau,\Delta}^{\mathrm{state}}(\widehat s) :=
I
\!\left(
Y_{i,t}^{(\tau,\Delta)};
\mathsf I_{i,t}
\mid
\widehat s_{i,t},X_t
\right).
\]

Read it this way: after the model knows the operational state and the proposition, how much additional information about the future remains hidden in the richer ideal state? Zero means the compression was sufficient for this outcome under the declared observational regime.

A fitted predictor may still misuse a sufficient state. Define the model-estimation gap

\[
\epsilon_{\theta,\tau,\Delta}^{\mathrm{model}}(\widehat s) :=
\mathbb E
\!\left[
D_{\mathrm{KL}}
\!\left(
\mathbb P(Y\in\cdot\mid\widehat s_{i,t},X_t)
\;\middle\|\;
P_{Y,\theta,\tau,\Delta}(\cdot\mid\widehat s_{i,t},X_t)
\right)
\right].
\]

The first error asks whether the state threw useful information away; the second asks whether the predictor used the retained information correctly. “The model was bad” is too imprecise to be useful, because better training cannot recover information the state discarded, while a richer state does nothing if the predictive head cannot use it.

<a id="gids-e14"></a>
**GIDS–14 — State error plus model error.**

Under ordinary log-loss regularity,

\[
\mathcal R_{\log}(P_{Y,\theta}\circ\widehat s)
-
\mathcal R_{\log}^{\star} =
\epsilon_{\tau,\Delta}^{\mathrm{state}}(\widehat s)
+
\epsilon_{\theta,\tau,\Delta}^{\mathrm{model}}(\widehat s).
\]

This equation earns its place because it tells us where to look. It is related to the information-bottleneck idea of compressing one variable while preserving what matters for another; GIDS asks for a broader preservation problem across a family of proposition-conditioned futures rather than one target alone.[^ib]

## Memory Is a Field of Weighted Traces

Memory need not begin as narrative. For the model, it may begin as a field of traces with changing weights:

\[
\mathbf m_{i,t}^{\mathrm{mem}} =
\sum_{j=1}^{N_i}
\varpi_{ij,t}\mathbf h_{ij}^{\mathrm{mem}}.
\]

Each trace representation carries a present availability or force; some decay, some repeat until they become structure, and some remain dormant until a proposition resembles the original event. A proposition-conditioned retrieval rule changes those weights,

\[
\widetilde\varpi_{ij,t}(x_t) =
\mathcal R_{\mathrm{ret},\theta}
\!\left(
\varpi_{ij,t},
\mathbf h_{ij}^{\mathrm{mem}},
\widehat s_{i,t},
 x_t
\right),
\]

then forms the local retrieved past

\[
\widetilde{\mathbf m}_{i,t}^{\mathrm{mem}}(x_t) =
\sum_j
\widetilde\varpi_{ij,t}(x_t)\mathbf h_{ij}^{\mathrm{mem}}.
\]

The present proposition does not consult the whole archive evenly; it retrieves a local past, and after the actual trace arrives the memory field changes again. The second encounter is therefore never with exactly the same actor-state as the first, because the first encounter has joined the actor.

A recommender system supplies the crude engineering analogy. A view history is not a mind, yet it demonstrates that repeated traces can be compressed into a latent object that improves prediction; GIDS takes the move seriously enough to separate source, role, memory, durable actor structure, and proposition-conditioned retrieval.

## Categorical Traces and the Registry

A great deal of useful evidence arrives categorically: roles, recurring topics, objection families, action types, counterpart identities, product themes, price postures, institutional regimes, and source channels. Before pooling, the model should lift surface labels through context,

\[
\widetilde{\mathcal B}_{i,r}^{(f,\sigma)} =
\operatorname{Lift}_{\mathrm{ctx}}
\!\left(
\mathcal B_{i,r}^{(f,\sigma)},
 c_{i,r}
\right),
\]

because “aggressive” toward a competitor, “deferential” toward a regulator, and “protective” toward a subordinate may express one deeper organization under different relations, while collapsing them at ingestion manufactures contradiction out of context.

Source remains explicit. Biography, stated language, observed behavior, and third-party inference are not merged merely because they share a label; a person may describe themselves as cautious, behave recklessly, and be described by others as calculating, and the disagreement is evidence. Slow categorical memory pools durable evidence by role and regime; fast retrieval emphasizes recent proposition-relevant traces; weighting may depend on recency, repetition, source reliability, action intensity, regime similarity, and estimated susceptibility. Missing cells use learned null representations and explicit masks, while count or evidence-mass terms keep one exposure from becoming numerically identical to twenty repeated exposures.[^categorical-pooling]

The registry behind the machinery is the current typed hypothesis over families, sources, regimes, interactions, masks, and reparameterizable latent blocks. It is not the species projection introduced earlier; it is the finite catalog of distinctions the current model knows how to ask about, revised under cross-task transfer. The strongest discovered content remains deliberately vague, while the method, chronology, and test conditions remain public.

## Identifiability and the Right Kind of Modesty

A learned actor-state can be rotated, rescaled, or reparameterized while preserving every prediction, which is not fatal; it means the empirical target is stable predictive information rather than one blessed coordinate chart. The slow/fast decomposition is more substantive because it predicts different failure patterns: slow state should help under sparse observation, role transfer, and longer horizons, while fast state should help after recent events and over shorter horizons; if both disappear without consequence, the decomposition has failed for that actor class.

The standard is not metaphysical proof. A distinction must persist where it should persist, change where it should change, transfer where it claims to transfer, and—when intervention data exists—participate in the predicted change under a proposition. At this point the actor can be written at several resolutions, from inherited seed to realized actor, phenomenal state, general predictive state, and operational estimate; the next move is to construct actors larger than one person without committing the obvious sin of averaging everyone together.

---

[^recap-operational-state]: This is [GIDS–11 — Slow and fast operational state](03_The_Chimera.md#gids-e11), repeated because the predictive-state chapter must be locally readable; no new mathematical claim is introduced.
[^psr]: Michael L. Littman, Richard S. Sutton, and Satinder P. Singh, “Predictive Representations of State,” *Advances in Neural Information Processing Systems 14* (2001), 1555–1561; see also Satinder Singh, Michael R. James, and Matthew R. Rudary, “Predictive State Representations: A New Theory for Modeling Dynamical Systems,” *Proceedings of UAI 2004*, 512–519.
[^ib]: Naftali Tishby, Fernando C. Pereira, and William Bialek, “The Information Bottleneck Method,” *Proceedings of the 37th Annual Allerton Conference on Communication, Control, and Computing* (1999); arXiv:physics/0004057.
[^categorical-pooling]: A compact implementation keeps family and source typed. For lifted bag \(\widetilde{\mathcal B}_{i,r}^{(f,\sigma)}\), use \(\mathbf u_{i,r}^{(f,\sigma)}=|\widetilde{\mathcal B}|^{-1}\sum_{\upsilon\in\widetilde{\mathcal B}}E_{f,\sigma}(\upsilon)\) when nonempty and a learned \(\mathbf e_{\varnothing}^{(f,\sigma)}\) otherwise; concatenate an aligned representation with an availability mask and \(\log(1+|\widetilde{\mathcal B}|)\). Slow banks pool these cells separately by role or regime with durable weights \(\beta\); fast retrieval uses task-conditioned weights \(\alpha\) and retains total relevance mass. This preserves source, missingness, repetition, and regime without promoting the pooling arithmetic into a central theorem.
