15 problems
- 0 votes0 replies1 view
Transformer-based circumvention of the curse of dimensionality
The transformer architecture is used to estimate a context function , with smoothness index , and can represent a Nadaraya--Watson estimator. If the data…
- 0 votes0 replies1 view
Symbolic pattern emergence in LAWS expert classes
Let be a sufficiently capable base model trained on a corpus containing code, mathematics, and structured data, and let be a PLT trie node with probability…
- 0 votes0 replies1 view
The separation conjecture
– separation conjecture. . Under this widely believed conjecture, chain-of-thought computation would extend the expres…
- 0 votes0 replies0 views
The robust predicate–attention conjecture for LEGO length generalization
Let be the length of a LEGO task, whose input consists of predicate and answer clauses, and let attention patterns describe which clauses a transformer uses when predicting the…
- 0 votes0 replies0 views
The BCJRFormer final-token conjecture on synchronization-loss severity
The BCJRFormer model processes received sequences of length corresponding to codewords of length , and uses attention across sequence tokens. Final-…
- 0 votes0 replies0 views
The attention-head information conjecture for salient network structure
The attention heads in a transformer network carry information about relationships among input tokens. Attention-head information conjecture. Exploiting information carried by the…
- 0 votes0 replies0 views
The error-ratio conjecture for Varshamov–Tenengolts code correction
Let a codeword have a given length, and let errors be introduced into it before correction. The error-ratio conjecture. The ratio of errors in the codeword is more closely related…
- 0 votes0 replies0 views
Conjecture on the separation condition for the limiting support
Let be the limiting parameter distribution at time , and suppose that the origin is an interior point of…
- 0 votes0 replies0 views
The computational-power conjecture for prompt engineering in transformers
A transformer is a neural architecture whose input can include an engineered prompt supplying additional information or memory. Prompt-engineering computational-power conjecture. P…
- 0 votes0 replies0 views
The RASP length-generalization conjecture for transformers
RASP length-generalization conjecture. Based on extensive empirical results, transformers tend to length-generalize on tasks that can be solved by RASP.
- 0 votes0 replies0 views
Higher-order transformer-based integrators without increased cost
Let transformer-based neural networks be used as integrators for ordinary differential equations, in analogy with multi-step methods that can achieve arbitrary order of approximati…
- 0 votes0 replies0 views
The model-disorientation conjecture for randomly initialized models
Model-disorientation conjecture. Model disorientation leads randomly initialized models not to achieve their full potential, regardless of model size.
- 0 votes0 replies1 view
The conjecture on self-attention after cross-attention in GNOT
The GNOT attention block consists of a cross-attention layer followed by a self-attention layer; query points and input functions provide the information processed by these layers.…
- 0 votes0 replies0 views
Transformer layers as rate-reduction optimization schemes
Transformer rate-reduction conjecture. Layers of the Transformer emulate a more general family of gradient-based iterative schemes that optimize the rate reduction of all input tok…
- 0 votes0 replies0 views
Conjecture on the success of diagonal initialization for Burgers equation transformers
The input consists of initial conditions, and the target consists of solutions at ; these are represented by the blue and green curves, respectively, in Figure. The model uses…