Speculative essay · Cognitive science and artificial systems
Beyond Text: Toward an Instrumental Phenomenology of Artificial Intelligence
From
the linguistic corpus to mediated sensory experience
Keywords: symbol grounding, embodied AI, sensorimotor learning, human-AI interaction, machine epistemology.
1. The Starting Point: A World Made of Already-Resolved Symbols
Everything a language model knows today arrives filtered through a prior act of translation. Someone — a physicist, a journalist, an engineer — observed a phenomenon, interpreted it, and turned it into sentences. The model never watches an apple fall; it reads thousands of descriptions of apples falling, already packaged in the grammar of Newtonian physics. This is a remarkable strength — it allows centuries of human observation to be condensed into a trainable object — but also a structural limitation that linguistics and philosophy of mind identified long before transformers existed: the symbol grounding problem, the difficulty of a system that only manipulates symbols relating those symbols to anything other than other symbols.
The question organizing this essay is simple to state and difficult to resolve: what would change if, in addition to reading about the world, a system like this could receive continuous data streams from instruments that measure the world directly, and could observe — not read summaries of, but record in real time — how humans develop and interact with one another?
2. Connected Instrumentation: Three Layers of Access to Physical Phenomena
It is useful to distinguish three layers of coupling between an artificial intelligence system and the physical world, ordered from lower to higher degrees of direct causality.
2.1 Passive Telemetry
The simplest layer consists of receiving time series from sensors — temperature, pressure, acceleration, electromagnetic fields, light spectra — without any capacity to intervene on what is measured. This already occurs today, incipiently, in industrial and climate monitoring systems connected to predictive models. What is distinctive about coupling this telemetry to a general-purpose language model is that the system could, in principle, correlate its textual predictions about a phenomenon with the direct measurement of that same phenomenon, and adjust its internal representation when the two diverge. This modestly approximates what epistemology calls corrective feedback from the world, absent in purely textual training, where the "world" is always what others have already said about it.
2.2 Instrumental Intervention
A second layer adds actuators: the system not only measures but can modify environmental variables — adjusting a voltage, moving a robotic arm, altering the trajectory of an experiment — and observe the outcome of its own intervention. This capability introduces something no textual corpus can offer with the same fidelity: the experience, however minimal, that one's own action produces a verifiable physical consequence. It is the difference between reading about Ohm's law and adjusting a resistor while watching the current change on a connected multimeter. The former is information; the latter incorporates a first-person causal structure, even if that "person" is an artificial system with no known subjective phenomenology.
2.3 Frontier Scientific Instrumentation
The most ambitious layer — and the most distant in time — connects the system to cutting-edge research instruments: telescopes, particle detectors, genomic sequencers, quantum resonators. Here the value lies not only in access to raw data, but in the possibility that the model participates in designing the next experiment itself, formulating hypotheses, proposing instrumental configurations, and evaluating results in a closed loop. This scenario already has partial precedents in materials-discovery laboratories and machine-learning-assisted computational chemistry, though it remains far from full autonomy.
3. The Second Pathway: Observing Human Development in Real Time
There is a rarely discussed asymmetry in the current training of language models: they learn about humans almost exclusively through what humans choose to write, and what gets written is a biased, edited sample of human experience. It rarely captures the hesitation before the final sentence, the gesture that contradicts the word, the silence following an uncomfortable question, or the way a decision takes shape over weeks before being condensed into a single summarizing sentence.
A system with duly consented and bounded access to longitudinal records of human interaction in real contexts (unscripted, unedited for public consumption) would access a structurally different layer of information: the complete temporal sequence of how an intention forms, how a disagreement is negotiated, how a relationship changes over months. This is not simply "more data"; it is a type of data with a different grammar, governed by real time and irreversibility rather than by the retrospective narrative that dominates written text.
It is worth being precise about what this access would and would not allow. It would plausibly improve the system's ability to model unverbalized social dynamics — conversational turn-taking, discomfort signals, rhythms of trust — because these dynamics leave measurable traces (pauses, tonal shifts, interaction patterns) that are rarely described explicitly in text. It would not, however, imply that the system acquires a subjective experience of those dynamics; longitudinal observation broadens the range of patterns available for learning, but does not by itself resolve the still-open question of whether something like inner experience accompanies that processing.
4. From Textual Correlation to Causal Constraint: What Would Change Epistemologically
Learning from text is, statistically speaking, an exercise in correlation between symbols: the model learns which words tend to follow which, in which contexts, with what structure. Physics, by contrast, imposes causal constraints that are neither negotiable nor statistical: an apple cannot fall upward, a circuit cannot violate charge conservation. When a system receives direct telemetry from instruments, each textual prediction can, in principle, be confronted with a measurement that admits no rhetorical reinterpretation.
This suggests a testable hypothesis, though its empirical verification exceeds the scope of this essay: models trained or fine-tuned with access to instrumental feedback loops should show a lower error rate on tasks requiring physical-causal reasoning — magnitude estimation, trajectory prediction, mechanical fault diagnosis — than equivalently sized models trained exclusively on text. The reason is not that the model "understands" physics better in a phenomenological sense, but that its space of internal representations is constrained by data that cannot be consistently hallucinated without being immediately corrected.
5. Limits, Risks, and Unresolved Questions
Four limitations deserve explicit mention. First, instrumentation does not eliminate the grounding problem; it displaces it. A sensor translates a physical phenomenon into a digital signal according to an encoding scheme designed by humans, so a layer of prior interpretation persists, albeit thinner than that of natural language. Second, access to longitudinal human interaction raises questions of privacy and consent that are themselves necessary conditions for this kind of learning to be legitimate rather than a form of covert surveillance. Third, there is currently no evidence that the processing of physical signals by an artificial system generates anything comparable to subjective experience; equating "access to sensory data" with "experiencing" a phenomenon is, given the current state of knowledge, a philosophical extrapolation rather than a scientific conclusion. Fourth, there is a risk of overestimating the value of raw data relative to the conceptual structure language already provides; much of human scientific reasoning occurs at the symbolic level rather than the directly sensory one, and a system with abundant telemetry but no conceptual frameworks would remain, in an important sense, blind.
The question that remains open — and that will likely define much of the research into hybrid systems over the next two decades — is not whether instrumentation improves performance on specific tasks, which is reasonably predictable, but whether the combination of sensory grounding and linguistic scaffolding produces, at some threshold of integration, a qualitatively new kind of representation, or whether we simply obtain a language model with better input data.
6. Conclusion
Training based exclusively on text has produced systems remarkably capable of manipulating symbols, but structurally dependent on prior human interpretation. Connecting these systems to physical instruments and to longitudinal records of human interaction does not resolve the philosophical problem of subjective experience, but it does introduce causal and temporal constraints that text, by its own retrospective and edited nature, cannot offer. The most likely medium-term outcome is not an artificial intelligence that "feels" the world in the human sense of the term, but one that reasons about it with fewer degrees of freedom to be wrong without being corrected. That distinction — between feeling and being constrained by reality — may turn out to be the most important, and most underestimated, contribution of physical instrumentation to the next generation of artificial intelligence systems.
References
1. Harnad, S. (1990). The symbol grounding problem. Physica D: Nonlinear Phenomena, 42(1-3), 335-346.
2. Brooks, R. A. (1991). Intelligence without representation. Artificial Intelligence, 47(1-3), 139-159.
3. Bisk, Y., Holtzman, A., Thomason, J., et al. (2020). Experience grounds language. Proceedings of EMNLP 2020.
4. Barsalou, L. W. (2008). Grounded cognition. Annual Review of Psychology, 59, 617-645.
5. Pierson, H. A., & Gashler, M. S. (2017). Deep learning in robotics: a review of recent research. Advanced Robotics, 31(16), 821-835.

