If you see this, something is wrong
First published on Thursday, Sep 24, 2026 and last modified on Thursday, Sep 24, 2026 by François Chaplais.
Orignal article author
Original article author
Talking with Caude Haiku in Documents Live
This is the transcript of a discussion between François Chaplais (WebMagic, creator of Documents Live) and Claude’s Haiku within Documents Live. A document set was created with A Philosophical Introduction to Language Models – Part I: Continuity With Classic Debates and A Philosophical Introduction to Language Models - Part II: The Way Forward . Then I had a chat with Claude’s Haiku within Documents Live, based on this document set.
Note: I had this discussion before on my development machine and the answers were quite different. I do not find this shoking: the questions are quite open.
The transcript follows.
Part 1 — Summary of Historical Foundations of LLMs
The historical development of Large Language Models can be traced through several interconnected intellectual traditions that ultimately converged to create modern approaches.
The Symbolic-Stochastic Divide
Early Natural Language Processing was characterized by a fundamental schism between two competing paradigms. The symbolic approach, influenced by Noam Chomsky’s transformational-generative grammar, sought to capture language through formal rules that could generate well-formed sentences. This led to rule-based syntactic parsers used in early conversational systems. In parallel, the stochastic approach, pioneered by Warren Weaver and influenced by Claude Shannon’s information theory, proposed using statistical techniques and computers for language tasks like machine translation. Initially, the symbolic paradigm dominated, while stochastic methods showed only modest success.
The Distributional Hypothesis
A crucial theoretical turning point came with linguist Zellig Harris’s distributional hypothesis in the 1950s, which posited that words acquire meaning through their patterns of co-occurrence with other units. This idea—crystallized in J.R. Firth’s observation that “You shall know a word by the company it keeps”—suggested that meaning emerges from context rather than being inherent to words themselves. This concept drew on Ludwig Wittgenstein’s philosophy of meaning-as-use.
From Multidimensional Spaces to Neural Embeddings
Building on the distributional hypothesis, researchers began exploring the possibility of representing word meanings as vectors in multidimensional space . Early psychological work rated words along dimensions like valence and potency, while later information retrieval research developed automated, data-driven techniques for vector-based representations. This culminated in neural network-based word embedding models that could learn distributional properties by predicting word contexts. Models like Word2Vec demonstrated that embeddings could capture semantic and syntactic regularities—famously, simple vector arithmetic revealed relationships like \( king + woman - man \approx queen\) .
Limitations and Evolution
However, word embeddings had critical limitations: they assigned static representations to each word (failing to capture polysemy), relied on shallow neural architectures with single hidden layers, and treated language as a “bag of words,” losing compositional structure and word order. These shortcomings led to deeper architectures—recurrent neural networks (RNNs), long short-term memory (LSTM) networks, and gated recurrent units (GRU)—which incorporated memory mechanisms for processing sequential information. The breakthrough came with the Transformer architecture , which laid the groundwork for modern LLMs by moving beyond sequential processing limitations.
Further Reading
The Mathematical Process Behind Modern LLMs
Modern Large Language Models operate through a sophisticated pipeline of mathematical transformations, centered on the Transformer architecture and the learning objective of next-token prediction.
The Input Pipeline: Tokenization and Embedding
The process begins with tokenization, where raw text is converted into discrete units. Rather than operating on whole words, modern LLMs break text into tokens that can map onto words or subword units . This choice reflects a computational trade-off: keeping the vocabulary size manageable while preserving linguistic expressivity. This quirk of tokenization has practical consequences—for instance, GPT-3’s tokenizer handles the number “940” as a single token but splits “941” into two tokens (“9” and “41”), which partially explains why LLMs can struggle with multi-digit arithmetic.
Each token is then converted into a vector representation through an embedding matrix. This transforms discrete tokens into points in high-dimensional continuous space, where semantically and syntactically similar tokens occupy nearby regions. Crucially, positional encoding is added to these embeddings to preserve information about each token’s location in the sequence, since the subsequent self-attention mechanism processes all tokens in parallel.
The Core Mechanism: Self-Attention
At the heart of modern LLMs lies the self-attention mechanism . This mechanism enables each token in a sequence to interact with every other token, computing attention scores that quantify the relevance or importance of each token to every other token.
Mathematically, for each token \( t_i\) in the sequence, the attention mechanism assigns a weight or “attention score” to every token \( t_{0-i}\) up to and including \( t_i\) . These scores reflect how much the model should “attend to” or consider each part of the input when processing a given token. For example, when processing the pronoun “it,” the attention mechanism determines which previous words “it” most likely refers to—allowing the model to resolve pronouns correctly even when they appear multiple times in different sentences with different antecedents.
The Transformer architecture employs multiple attention heads , where each head specializes in attending to different types of dependencies. One head might focus on syntactic relationships, another on long-range semantic connections, and another on discourse structure. This parallel processing of diverse dependency types dramatically enhances the model’s ability to capture complex linguistic patterns.
A critical architectural advantage of the Transformer is parallel processing: all words in the input sequence are processed simultaneously rather than sequentially (as with RNNs, LSTMs, and GRUs). This parallelization dramatically improves training efficiency and enables the model to handle longer sequences more effectively. However, this comes with a computational constraint: self-attention’s memory and computational requirements scale quadratically with sequence length, limiting how long the input window can be.
The Transformer Block: Attention and Feed-Forward Layers
Each Transformer block consists of two main components:
These layers are stacked sequentially, with the output of one block feeding into the input of the next. Modern LLMs typically stack dozens to hundreds of these blocks, allowing information to be progressively refined through multiple levels of abstraction.
The Output Stage: Unembedding and Next-Token Prediction
After processing through all Transformer blocks, the token vectors undergo a linear transformation called unembedding that projects them into a vocabulary-sized space. This produces a set of raw output scores called logits—one logit for each possible token in the model’s vocabulary. Each logit represents the model’s preliminary assessment of how likely that token is to be the next element in the sequence, given the input.
A softmax function is then applied to convert these logits into a probability distribution over the vocabulary, yielding normalized probabilities that sum to one. During training, the model compares its predicted probability distribution against the actual next token in the training data, computing a loss (typically cross-entropy loss) that measures the discrepancy between prediction and ground truth.
The Learning Objective: Next-Token Prediction
Modern LLMs are trained using autoregressive next-token prediction. Given a sequence of tokens, the model learns to predict which token is statistically most likely to follow. During training:
This learning objective is remarkably simple yet powerful: by optimizing for next-token prediction across an enormous diversity of contexts, the model implicitly learns rich representations of language structure, semantics, and world knowledge.
Post-Training Refinement: RLHF
While next-token prediction produces models excellent at generating fluent text, it has no built-in preference for truthfulness, usefulness, or harmlessness. To address this, modern LLMs like ChatGPT undergo post-training refinement using Reinforcement Learning from Human Feedback (RLHF), which proceeds in three stages:
This iterative process allows developers to steer model outputs toward specific normative goals while preserving the model’s generative capabilities.
In-Context Learning and Inference
Deployed LLMs operate with frozen parameters—they do not learn in the conventional sense during text generation. However, they demonstrate a remarkable capacity called in-context learning: they flexibly adjust their outputs based on contextual information provided in the prompt, including for tasks they have not been explicitly trained on.
In few-shot learning, the prompt includes a few examples of a task followed by a new instance requiring response. The model, aiming to complete the pattern, produces output consistent with the examples. In zero-shot learning, no examples are provided; instead, the task is stated directly as an instruction. Modern LLMs—particularly those fine-tuned with RLHF—excel at both paradigms, leveraging their extensive training exposure to parse instructions and generate appropriate responses.
Further Reading
LLMs in Context: Parametric Modeling, Dynamical Systems, and Functional Identification
Your background in parametric dynamical systems and functional identification via splines provides valuable conceptual scaffolding for understanding LLMs, though the relationship is more subtle than direct analogy.
Parametric Modeling and the Architecture Analogy
In dynamical systems, you work with explicit parametric forms—differential equations with tunable coefficients that govern how a system evolves through state space. LLMs similarly employ parametric models: the Transformer architecture defines a fixed computational structure, and training tunes millions to billions of numerical parameters (weights and biases) to optimize next-token prediction. Both are therefore parametric systems where learning amounts to parameter optimization.
However, a critical disanalogy emerges: dynamical systems typically assume an underlying generative mechanism—you hypothesize a structural form (e.g., a system of ODEs) based on domain knowledge, then fit parameters to observed trajectories. LLMs, by contrast, employ an architecture-agnostic learning objective: next-token prediction. The Transformer’s structure is engineered rather than derived from first principles about language or cognition. Parameters are optimized not to match a theoretical model but to minimize prediction error on training data.
This echoes a key tension in your work: spline-based functional identification often involves choosing basis functions (knots, degrees) without knowing a priori what functional form best represents the underlying dynamics. LLMs face an analogous problem—the architecture is largely a design choice, not a consequence of reasoning about language structure.
Function Approximation and Neural Networks as Universal Approximators
Your experience with splines as flexible function approximators connects directly to a foundational result in deep learning: the universal approximation theorem. Just as splines can approximate arbitrary smooth functions to arbitrary precision (given sufficient basis functions and knots), sufficiently wide neural networks can approximate any continuous function on compact domains.
However, this connection masks a deeper divergence:
This raises a central interpretability challenge: your spline-based methods produce identifiable functions—you can recover the basis representation and understand the mapping. LLMs produce opaque parameter matrices whose internal organization remains largely mysterious.
State Space and Representation Geometry
A productive connection emerges at the level of representation learning. In dynamical systems, the phase space structure encodes the system’s essential behavior: stable manifolds, attractors, bifurcation geometry. Similarly, LLMs learn high-dimensional vector spaces—token embeddings, intermediate activations in the residual stream —where meaningful linguistic and semantic structure emerges geometrically.
For instance, word embeddings exhibit the property that semantically and syntactically similar tokens occupy nearby regions in representation space, much as trajectories in phase space cluster near stable manifolds. Interpretability research has discovered that LLMs organize representations along interpretable dimensions—gender, number, tense—corresponding roughly to coordinate directions in embedding space, analogous to slow-fast decomposition or other coordinate-theoretic structures in dynamical systems.
The residual stream is particularly suggestive of dynamical thinking. Information flows through the network layer-by-layer, with each Transformer block reading from and writing to this central vector. This resembles a discrete-time dynamical system: \( x_{t+1} = x_t + f_{\theta}(x_t)\) , where \( f_{\theta}\) represents the combined action of attention heads and MLPs in a single layer, and \( x_t\) is the residual stream at layer \( t\) . The residual connection architecture is actually motivated by deep learning work on training stability, but it can be reframed as implementing a recurrent structure akin to iterative refinement in dynamical evolution.
Mechanistic Interpretability as Functional Identification
Your concern with functional identification—recovering the underlying mapping from observed input-output behavior—parallels emerging work in mechanistic interpretability. This field aims to decompose LLM computations into understandable subcomponents, much as you might decompose a complex dynamic into simpler functional modules.
A key result involves identifying circuits: specific configurations of interconnected weights that collectively perform a computation. For example, researchers have isolated induction head circuits that implement a non-content-specific algorithm: find previous token occurrences matching the current token, attend to what followed them, and predict that same continuation. This is functional identification in your sense—recovering the algorithmic mechanism underlying observable behavior.
Similarly, causal intervention methods directly probe representations by manipulating them (via techniques like interchange interventions) and observing effects on outputs. This mirrors experimental perturbation approaches in dynamical systems: apply a controlled stimulus, measure response, infer underlying structure.
Efficiency, Abstraction, and Generalization
A crucial limitation of LLMs relative to human cognition, noted in the retrieved literature, is learning efficiency. You work with parametric models partly because they leverage domain structure to achieve efficiency—far fewer parameters and training examples than required to fit an unstructured neural network to the same data.
The retrieved text emphasizes that current LLMs implement important aspects of abstractness, systematicity, and generalizability characteristic of human cognition, but fall short in efficiency, completeness, and agency . They require orders of magnitude more training data than humans to achieve comparable performance—precisely because they lack the structured architectural priors (modular organization, domain-specific processing) that humans possess.
This suggests a research direction: incorporating insights from parametric dynamical systems—explicit basis decomposition, sparse representation, modular structure—might improve LLM efficiency. Rather than relying purely on end-to-end gradient descent on monolithic parameter matrices, hybrid approaches might impose structured priors that make the learned representations more interpretable and efficient.
Splines, Basis Functions, and Distributed Representations
Spline-based methods excel partly because they employ interpretable basis functions with clear geometric meaning (polynomials of specified degree on specified intervals). LLM parameters, by contrast, are distributed across the entire network—a single neuron contributes to many computations, and a single computation involves many neurons. This distributed representation is both a strength (robustness via redundancy, graceful degradation) and a weakness (opacity, difficulty in extraction and interpretation).
Interestingly, mechanistic interpretability research has identified structures resembling basis-like organization: attention heads that specialize in specific linguistic tasks (detecting syntax, tracking discourse, etc.), MLP layers that appear to implement lookup tables or classification operations. However, these bases are learned rather than imposed, and their interpretation remains contested.
A speculative direction: could structured spline-like basis functions be incorporated into LLM architectures explicitly? For example, replacing fully-connected MLP layers with basis-function expansions, or decomposing attention patterns via learned spline dictionaries? This could improve interpretability and efficiency at the potential cost of architectural simplicity and empirical performance.
Identifiability and the Interpretation Challenge
In your parametric dynamical systems work, identifiability is a central concern: given observed trajectories, is the parameter set uniquely recoverable? LLMs face an analogous but more severe problem: even when representations are known to encode specific linguistic properties (e.g., via probing), it remains unclear whether the model actually uses that information causally in its computations . Probing reveals correlations; causal interventions are required to establish functional relevance.
Moreover, interpretations of LLM representations are often non-unique—alternative representational structures may fit the same behavioral data equally well , raising fundamental questions about whether networks implement determinate computational strategies or merely approximate input-output mappings without internal structure.
Further Reading
Mallat’s Wavelet Bases to Deep Learning: Continuity in Representation Learning
Your intuition about continuity in Mallat’s work is profound and largely correct, though the connection involves some subtle shifts in philosophical orientation. Let me trace the intellectual lineage and then locate LLMs within this broader framework.
The Core Thread: Efficient Representation and Sparse Coding
Mallat’s foundational insight was that meaningful structure in signals emerges through sparse representation in appropriate bases. His work on wavelet decomposition, multiscale analysis, and best basis search established that:
The wavelet framework provided explicit, interpretable basis functions (Daubechies wavelets, Mexican hats, etc.) with clear geometric and smoothness properties. Best basis search algorithms systematically explored candidate bases to find those yielding maximum sparsity—a principled approach to representation optimization.
The Transition to Deep Learning and Scattering Networks
Mallat’s transition to deep learning preserved this core philosophy but adapted it to higher-dimensional, more complex domains. Rather than hand-crafted wavelets, deep networks learn representations end-to-end. Yet the underlying principle persists: successive layers learn increasingly abstract, sparse features that compress information hierarchically.
His scattering networks formalize this intuition: a fixed multilayer architecture (cascade of wavelet transforms with nonlinearities) that provably extracts stable, invariant features without learning. Scattering networks avoid the black-box quality of fully-learned deep networks while preserving the hierarchical feature extraction principle. They represent a middle ground between engineered wavelets (interpretable but inflexible) and learned networks (flexible but opaque).
Connecting to LLMs: The Abstraction Hierarchy
LLMs exhibit a structure remarkably similar to wavelet-based hierarchical decomposition, though the basis is learned rather than engineered:
Token embeddings (shallow layers) capture raw linguistic tokens as dense vectors in a learned metric space, analogous to the finest wavelet scale capturing high-frequency details.
Intermediate layers progressively compose embeddings, building representations of word sequences, phrases, and semantic relationships—analogous to coarser wavelet scales that aggregate information while preserving structure.
Deeper layers extract increasingly abstract features: syntactic patterns, discourse dependencies, world knowledge, reasoning operations—analogous to the coarsest scales in a multiscale decomposition that capture the global signal structure.
This hierarchical abstraction is not incidental to LLMs; it’s central to their function. When you apply mechanistic interpretability techniques , you discover that intermediate layers specialize in particular computations: early layers handle token and syntactic features, middle layers track semantic and discourse structure, and later layers perform high-level reasoning and planning. This division of labor across scales echoes Mallat’s multiscale decomposition philosophy.
Sparsity and Lottery Tickets
A particularly striking continuity emerges from recent work on neural network lottery tickets. Research has shown that most parameters in trained neural networks are redundant; sparse subnetworks (carefully pruned from the full network) can achieve comparable performance with far fewer parameters. This is precisely Mallat’s sparsity principle applied to learned networks: the “right” basis (in this case, the right subnetwork) is sparse and concentrated.
This connects directly to your background with spline-based functional identification. Just as you impose basis constraints to achieve identifiability and efficiency, lottery ticket work suggests that learned representations naturally concentrate in sparse configurations—networks effectively discover efficient, interpretable bases through training, even though we have not explicitly imposed sparsity constraints.
Interestingly, mechanistic interpretability studies of algorithmic learning show networks learning algorithms for modular arithmetic via a “Fourier multiplication” circuit. Rather than memorizing all input-output pairs, the network discovers a sparse, elegant algorithm—a basis-like representation of the computation. This is functional identification in the strongest sense: the network learns not just to predict outputs but to decompose the computation into understandable, reusable components.
Best Basis Search as Meta-Learning
A deeper connection involves best basis search itself. Mallat’s algorithms systematically explored a dictionary of candidate bases to find those yielding maximum sparsity for a given signal class. In modern machine learning terms, this is a form of meta-learning: learning about learning—discovering which representation is optimal for a task class.
LLMs engage in analogous meta-learning implicitly. During pretraining on next-token prediction across enormous, diverse text corpora, they implicitly learn which representational structures are generally useful. In-context learning—where LLMs adapt to new tasks from a few examples—represents this meta-learned representation being applied flexibly to novel scenarios.
More directly, recent work on learned optimizers and architecture search uses neural networks to discover good architectures and learning algorithms for specific domains—a computational realization of Mallat’s best basis philosophy. Rather than searching over a hand-crafted dictionary of wavelets, we search over network architectures and learning algorithms, using differentiable search methods to optimize for task performance.
The Interpretability Imperative
Here the continuity becomes subtle and philosophically interesting. Mallat’s wavelet work emphasized interpretability: the basis functions had clear geometric meaning, and sparse representations could be visualized and understood. Scattering networks preserved this emphasis despite learning features—the architecture itself provides interpretability structure.
LLMs, by contrast, learn distributed, opaque representations. Yet mechanistic interpretability research increasingly reveals that learned LLM representations exhibit structure remarkably similar to engineered bases. Attention heads specialize like wavelet filters. MLP layers implement lookup tables or nonlinear transformations resembling wavelet nonlinearities. Residual streams accumulate and refine information hierarchically.
This suggests a unifying perspective: learned representations naturally discover interpretable structure when trained on appropriate objectives and architectures. The wavelet basis was not arbitrary; it emerged from principled reasoning about signal structure and smoothness. Similarly, LLM representations emerge from principled reasoning about language structure (via next-token prediction) and architectural constraints (attention, residual connections, layer normalization).
Mallat’s insight thus extends in both directions:
The Limitation: Efficiency and Data Hunger
However, the continuity is incomplete in one crucial respect. Wavelet-based methods achieved efficiency: a modest number of hand-crafted basis functions could represent images compactly, requiring minimal training data. LLMs require orders of magnitude more data and parameters to achieve comparable generalization to human cognition.
This echoes a theme in the retrieved literature: LLMs lack the efficiency that characterizes human learning . They implement abstraction and compositionality but wastefully, through sheer scale rather than principled architectural constraints.
This suggests a research direction fully in line with Mallat’s philosophy: incorporate structured priors into LLM architectures to improve efficiency. Rather than learning all parameters end-to-end, impose modular structure, sparse connectivity, or hierarchical organization inspired by wavelets and scattering networks. This could yield models that are simultaneously more interpretable and more efficient.
Vision-Language Models and Multimodal Basis Learning
Interestingly, recent advances in multimodal models and vision-language models embody a Mallat-like principle: learning shared representations across modalities (text and images) requires finding a common basis—a metric space where both visual and linguistic features can be meaningfully compared.
CLIP’s contrastive learning objective learns precisely this: a shared latent space where image vectors and text vectors of matching captions align. This is best basis search across modalities: finding the representational structure that most efficiently captures cross-modal relationships. Vision-language models build on this foundation, using language as a “universal API” (in the terminology of the retrieved literature) that bridges perception and generation.
Mallat’s wavelets decomposed images via multiscale basis functions. Modern multimodal models decompose perception and language into a shared abstract space. The philosophical continuity is clear: find the representation that most efficiently captures task-relevant structure.
Summary: The Unifying Principle
Mallat’s scientific trajectory reflects a single, evolving principle:
Efficient representation emerges through discovering sparse, hierarchical bases tailored to problem structure. These bases may be engineered (wavelets), partially learned (scattering networks), or fully learned (deep networks), but the underlying logic—that meaningful structure concentrates in low-dimensional subspaces—remains constant.
LLMs represent the full-learning extreme of this spectrum. They start from relatively unprincipled architectural choices (Transformers were engineered for computational efficiency, not derived from first principles about language) and discover through training that sparse, hierarchical, interpretable structures emerge. The continuity in Mallat’s work suggests this is no accident: when you optimize for appropriate objectives (next-token prediction, multimodal alignment) on diverse data, networks naturally learn representations that resemble hand-crafted bases—because such bases reflect genuine structure in data.
The frontier of research lies in bridging this gap: incorporating the interpretability and efficiency of engineered bases (wavelets, scattering networks) with the flexibility and power of learned representations, yielding systems that are simultaneously more capable and more understandable.
Further Reading
History Recompaction in LLMs: Projection, Compression, and Dynamical Systems
Your intuition about “recompacting history” as projection onto a smaller space is insightful and touches on several deep issues in how LLMs process sequential information. Let me unpack this carefully, as it connects to representation theory, information theory, and mechanistic interpretability.
What “Recompacting History” Means in LLMs
When we say an LLM “recompacts its history,” we typically mean one of two related phenomena:
Both phenomena can be framed as lossy compression: extracting a lower-dimensional summary of the past that retains task-relevant information while discarding task-irrelevant details.
The Projection Interpretation: Partially Correct
Your projection intuition is partially but not entirely correct. Here’s why:
Where Projection Applies
A projection onto a lower-dimensional subspace is mathematically precise when:
In some interpretability studies, this formalization is nearly exact. For instance, when researchers use iterative nullspace projection to remove information about linguistic features (like relative clause boundaries), they are literally projecting representations onto nullspaces—removing detectable information about a target feature while preserving other information. The mathematical operation is projection.
More broadly, since mechanistic interpretability reveals that LLM representations organize along interpretable dimensions , the idea that recompaction involves projection onto task-relevant subspaces is plausible. Early tokens’ fine-grained syntactic and semantic features might lie in high-frequency dimensions irrelevant to next-token prediction; recompaction could involve discarding those dimensions, retaining only dimensions encoding abstract discourse structure or semantic macrostate.
Where Projection Breaks Down
However, several complications arise:
1. Nonlinearity and Information Bottlenecks
The residual stream architecture processes information through multiple nonlinear layers. Information flow is not a single linear projection but a cascade of nonlinear transformations:
Each attention head and MLP layer can selectively amplify or suppress information about different aspects of history. This is selective filtering rather than orthogonal projection—information is processed through nonlinear gates and mixing operations that depend on content.
For example, an attention head might attend selectively to earlier tokens that are semantically or syntactically relevant to the current token, ignoring irrelevant distant context. This is more sophisticated than projection: it’s adaptive, content-dependent compression.
2. Lossy Compression via Information Bottleneck
From information theory, context compression in sequence models is often understood through the information bottleneck principle: as the model processes longer sequences, the hidden state is constrained to carry finite information (bounded by its dimensionality). Later states must compress earlier information, retaining what is predictive of future outputs while discarding what is not.
This is not orthogonal projection but lossy compression: different information is discarded depending on context. For example:
The same distant context (the bank, the robber) is compressed differently depending on what information proves predictive downstream. This is context-dependent projection, not linear projection onto a fixed subspace.
3. Superposition and Polysemy
Recall from the retrieved literature that LLMs can represent many more features than they have dimensions through superposition—overlapping distributed representations where the same neurons encode multiple features . When information is compressed, it’s not being orthogonally projected away but rather entangled with other information in a higher-dimensional code.
Recompaction thus involves not elimination but interference and entanglement: early token information becomes mixed with current information in a way that is difficult to linearly separate. This is qualitatively different from clean projection.
A Refined Framework: Adaptive Lossy Compression
A more accurate characterization of history recompaction combines ideas from information theory, dynamical systems, and representation learning:
The residual stream evolves as a discrete dynamical system (as you might frame it given your dynamical systems background):
where \( f_{\theta}\) represents the combined action of attention and MLP layers. This is additive update with content-dependent nonlinear feedback—each layer reads the state, applies a content-dependent transformation, and adds the result back.
This architecture naturally implements selective refinement and compression:
The result resembles a slowly-varying slow manifold in dynamical systems terminology: early token details occupy fast-varying, high-frequency components of the state that decay quickly, while abstract discourse structure occupies slow-varying, low-frequency components that persist.
This is not pure projection but manifold dynamics: information about history is not discarded orthogonally but rather pushed toward lower-dimensional slow manifolds that capture task-relevant structure while dissipating task-irrelevant details.
Quantifying Recompaction: Mutual Information
Information theory provides a more precise language. If we denote:
Empirical studies suggest something like:
where the exponential decay dominates over moderate distances, and \( \gamma\) represents a floor of coarse information about the distant past. This exponential decay mirrors information loss in many physical systems and biological memory (forgetting curves), and can arise naturally from the information bottleneck principle.
Equivalently, this could be framed as: at each layer, attention weights approximately implement a decaying projection that downweights distant tokens, with the decay controlled by task-dependent attention patterns.
Connection to Your Background: Phase Space Dynamics
Given your background with parametric dynamical systems, here’s a particularly relevant framing:
The residual stream at different layers can be viewed as a phase space trajectory in a high-dimensional space (the embedding dimension, typically 768–12,288 for modern LLMs). As the model processes a sequence:
This resembles center manifold reduction or slow-fast decomposition in dynamical systems:
The residual connection structure (additive updates) rather than replace operations facilitates this: at each layer, you don’t reset the state but rather refine it additively. This preserves slow-varying components while allowing fast components to dissipate through multiplicative gating (attention weights) and nonlinear filtering (MLP layers).
Grokking and Phase Transitions
Interestingly, the grokking phenomenon provides empirical evidence for sharp transitions in representation structure. During grokking:
This suggests that recompaction is not a passive degradation but an active learning process. The model discovers efficient, low-dimensional algorithmic representations through training, then progressively prunes away high-dimensional memorization noise.
This maps onto your spline-based functional identification intuition beautifully: just as you seek efficient parametric bases for dynamical systems, networks discover efficient algorithmic bases for language tasks. Grokking is the empirical manifestation of this discovery and cleanup.
A Speculative Connection: Wavelets and Recompaction
Returning to your earlier point about Mallat’s wavelets: multiscale wavelet decomposition naturally implements hierarchical recompaction. Fine scales (high frequency) capture detail; coarse scales (low frequency) capture structure. A wavelet-compressed representation discards fine-scale information while preserving coarse-scale information—precisely the recompaction you’re describing.
Modern LLMs might be implicitly implementing wavelet-like compression through their layer structure: each layer is a filtering operation that separates information into detail (high-frequency, task-irrelevant) and structure (low-frequency, task-relevant) components. The residual connection allows structure to flow through unchanged while detail is progressively filtered away.
If true, this suggests that incorporating explicit wavelet-like bases into LLM architectures might improve both interpretability (we’d understand the compression explicitly) and efficiency (we’d discard unnecessary information more systematically).
Summary
“Recompacting history” is partially but not fully captured by projection onto a smaller space. More precisely:
Your intuition about projection captures something real but incomplete—a good starting point for deeper formalization using information theory, manifold dynamics, and mechanistic interpretability.
Further Reading
Connecting LLMs and PINNs: Inductive Bias, Structure, and Physics-Informed Learning
Your question about connecting LLMs and PINNs (Physics-Informed Neural Networks) opens a fascinating dialogue between two very different paradigms for neural network design. This is particularly rich given your background with parametric dynamical systems, and it touches on fundamental issues about inductive bias, generalization, and the role of structure in learning.
The Fundamental Contrast: Inductive Bias vs. Flexibility
At the most basic level, LLMs and PINNs represent opposite design philosophies:
LLMs (Large Language Models):
PINNs (Physics-Informed Neural Networks):
This is precisely the tension you’ve been exploring with wavelets and parametric dynamical systems: Do you impose structure (wavelets, differential equations) and learn parameters, or do you learn structure and parameters end-to-end from data?
How PINNs Work: Physics as a Regularizer
To make this concrete, consider how a PINN operates. Suppose you want to learn a solution \( u(t, x)\) to a partial differential equation:
(This is the viscous Burgers equation, a canonical example.)
A traditional approach:
A PINN approach:
Fit a neural network \( u_\theta(t, x)\) by minimizing:
where:
The physics constraint is enforced via automatic differentiation: you compute derivatives of the network outputs with respect to inputs and check whether the differential equation is satisfied.
The Payoff of Physics-Informed Learning
PINNs achieve remarkable benefits from this structure:
This is precisely what you achieve with parametric dynamical systems: by imposing the structure of differential equations, you gain efficiency, interpretability, and ability to perform functional identification.
Can We Apply PINN Ideas to Language?
This is where the connection becomes subtle and speculative. The question is: Are there “physics equations” for language—structural principles as strong and universal as conservation laws—that we could embed into LLMs as inductive bias?
Candidate “Language Physics”
Several researchers have proposed principles that might serve this role:
1. Compositionality
The meaning of a complex expression should be computable from the meanings of its parts. This is a strong constraint: it implies that representations should be organized hierarchically, with lower layers capturing word meanings and higher layers combining them systematically.
Interestingly, LLMs do exhibit compositionality empirically, but they achieve it through training on diverse language, not through architectural constraint. A PINN-like approach might impose compositional structure explicitly—for example, by requiring that representations be linearly composable in some learned basis.
2. Syntactic Constraints
Natural language has grammatical structure: words have parts-of-speech, phrases have hierarchical constituency, sentences follow recursive rules. These are formal constraints, akin to conservation laws. One could imagine a PINN-like loss that penalizes violations of grammatical structure, enforced via a parser component.
3. Semantic Consistency
If a model asserts “The ball is red” and later “The ball is not red,” this violates a logical consistency constraint. Similarly, statements about the same entity should maintain consistent representations. One could add a loss term penalizing logical inconsistencies, detected via formal reasoning modules.
Barriers to Language PINNs
However, several fundamental obstacles prevent straightforward application of PINN ideas to LLMs:
1. Language “Physics” is Normative, Not Descriptive
Physics equations describe how systems must behave—they are discovered through observation and apply universally. Linguistic rules, by contrast, are largely conventions: English speakers typically follow grammatical rules, but we constantly violate them creatively. No equation describes this.
Moreover, linguistic rules are culture and history-dependent, learned through social interaction. There’s no universal “physics” of language like there is for fluids or elastic materials.
2. The Target Space is Discrete and Combinatorial
PINNs work on continuous state spaces where differential equations are naturally expressed. Language operates on discrete tokens and combinatorial structures. The appropriate mathematical framework might be formal language theory or logic rather than differential equations. It’s unclear how to express such constraints efficiently as neural network loss terms.
3. Measuring Constraint Violation is Non-Trivial
For Burgers equation, you can compute exact derivatives via autodiff and evaluate the residual precisely. For linguistic rules, detection and quantification are ambiguous. For instance, is “The book I was reading” grammatically well-formed? (Yes—relative clause attachment.) Is “Colorless green ideas sleep furiously” grammatical? (Syntactically yes; semantically dubious.)
4. The Cost-Benefit Calculus is Unfavorable
LLMs achieve remarkable generalization and compositionality without explicit linguistic structure constraints, simply by training on diverse text at scale. Adding explicit constraints might improve efficiency (requiring less data), but if you have access to internet-scale text corpora, the data efficiency gain may not justify the added complexity and reduced flexibility.
A More Feasible Direction: Hybrid Approaches and Modular Architectures
Rather than PINNs for language directly, a more promising direction involves incorporating PINN-like ideas in hybrid, modular systems where LLMs integrate with symbolic reasoning or physics-aware components.
Example 1: World Models with Physics Constraints
The retrieved literature discusses world models in LLMs and vision-language models . One could augment such systems with a PINN-like physics module:
The overall system would combine:
This resembles the modular architectures discussed in the retrieved literature , where language acts as a “universal API” between modules.
Example 2: Constrained Code Generation
An LLM fine-tuned to generate code (Python, C++, etc.) could be augmented with a PINN-like physics constraint checking that generated code respects:
The loss would combine:
This is closer to actual PINN philosophy: you have ground truth constraints (type systems, memory models) that should be strictly enforced.
Example 3: Scientific Discovery and Symbolic Regression
One of PINNs’ greatest strengths is discovering governing equations from data. A hybrid system combining LLMs with PINNs could:
This is a form of physics-informed language generation—constraining the LLM’s output space to hypotheses consistent with data and known physics.
Connecting to Your Background: Parametric Identification and Dynamical Systems
Here’s where your expertise becomes central. You work with parametric dynamical systems where:
PINNs are similar: they impose differential equation structure and learn parameters (network weights) to satisfy both data and physics.
LLMs take the opposite approach: they impose minimal structure and learn everything—architecture and parameters—end-to-end.
A natural middle ground would be:
Hybrid Parametric-Neural Networks: Systems where:
For example, you might have:
This is actively researched under names like “hybrid neural differential equations,” “universal differential equations,” and “gray-box modeling.” It combines the best of both worlds: structural efficiency from parametric models and flexibility from neural networks.
The Deep Connection: Inductive Bias and Generalization
Stepping back, the real insight connecting LLMs and PINNs is about inductive bias and generalization:
Strong inductive bias (as in PINNs):
Weak inductive bias (as in LLMs):
The question for your research might be: For language and cognition, what level of structural constraint is optimal?
Current LLMs suggest that very weak constraints + enormous data is sufficient to learn impressive language abilities. But they also exhibit profound inefficiencies and failure modes. Perhaps incorporating structured priors—inspired by PINNs but adapted to language and cognition—could yield systems that are simultaneously more capable and more efficient.
A Speculative Vision: Physics-Informed Language Models
Imagine a future architecture:
PILMs (Physics-Informed Language Models) would combine:
The entire system would be trained jointly on:
This is speculative, but it points toward systems that maintain LLMs’ flexibility while incorporating structured inductive biases that improve efficiency, interpretability, and robustness.
Connection to Mechanistic Interpretability
Interestingly, mechanistic interpretability research reveals that LLMs already implement many constraint-like structures implicitly . For example:
This suggests that implicit constraints emerge from appropriate training objectives and architectures. The PINN lesson—that explicit constraints improve efficiency—might be implemented not by adding penalty terms but by architectural innovations that naturally induce the desired inductive biases.
Summary
LLMs and PINNs are orthogonal approaches to a central challenge: how to incorporate structural knowledge into learned models.
The middle ground—Hybrid approaches incorporating both learned flexibility and structural constraints—may hold the key to more efficient, interpretable, and robust AI systems. For your work with parametric dynamical systems and functional identification, this suggests exploring:
Further Reading