NEWS 8 min read

PSSA Tests a Tiny Non-Transformer Language Model in Rust

PSSA reports a parameter-matched experiment in which a 1.5M-parameter recurrent state-space model beat a small transformer on held-out WikiText and generated faster on CPU. The repository is unusually candid about why that is not yet a general architecture victory.

By EgoistAI ·
PSSA Tests a Tiny Non-Transformer Language Model in Rust

An open-source project called PSSA is testing a language-model design that does not use transformer attention. Its author describes a plastic state-space architecture: a recurrent state-space core reads one token at a time, an episodic memory bank retrieves a small number of stored items, and fast-changing weights are periodically consolidated into the base model. The implementation is written from scratch in Rust rather than built on PyTorch or TensorFlow.

The headline numbers are interesting but deliberately narrow. The repository compares two models with about 1.5 million parameters, the same tokenizer, corpus, optimizer schedule, seed, and token budget. On an unseen 198,939-token WikiText-103 slice, PSSA reports cross-entropy of 3.997 versus 4.429 for its transformer baseline, with next-token accuracy of 24.1% versus 18.0%. In one same-CPU generation test, 200 tokens took 226 milliseconds for PSSA and 2,735 milliseconds for the transformer.

The Hacker News submission had 35 points and 9 comments when checked at 1:41 p.m. Malaysia time on September 30. That is a discovery signal, not peer review. The stronger evidence is the public code, documented comparison protocol, checkpoints, tests, and explicit list of unanswered experiments.

What happened

PSSA combines ideas that have appeared separately in recurrent models, state-space models, associative memory, and online learning. A fixed-size hidden state carries information forward, avoiding attention over every pair of tokens. A 512-slot memory bank supports bounded top-four retrieval. “Plastic” updates let part of the system adapt while it runs, while a closed-form ridge-regression step folds those updates back into a transition matrix.

The author trained PSSA and a small transformer over roughly 12.8 million tokens from a cleaned WikiText-103 stream. Both used the same 2,048-token vocabulary and matched parameter count. The repository reports that PSSA stayed ahead on the training curve and on a held-out slice. It also provides a scalar CPU reference and says CUDA gradients were checked against it, with a maximum difference of 2.98e-8.

That is a more useful announcement than an architecture diagram alone. The project states what was held constant, identifies where hardware was not matched, and includes commands intended to reproduce the runs.

Why it matters

Transformer attention is powerful, but its cost encourages research into architectures that carry state more efficiently. A recurrent model can process a sequence in one direction without repeatedly scanning the entire context. If a small system can preserve useful information through state and selective memory, it may offer practical advantages for local or CPU-bound inference.

The Rust implementation also matters. A framework-free reference makes numerical assumptions, memory layout, update rules, and inference costs more inspectable. Researchers can see whether an apparent speed gain comes from architecture, batching, kernels, or simply comparing different execution paths.

PSSA should therefore be read as an experiment about research directions, not a challenger to production foundation models. Its most valuable contribution today may be a compact test bed for asking which components—state, memory retrieval, plastic updates, or consolidation—actually create the reported gap.

Evidence

The best-supported claim is the small matched experiment. The repository documents model widths, vocabulary size, token counts, optimizer continuity, held-out evaluation, and the same-CPU generation test. It also publishes exact losses at multiple checkpoints instead of only the final number.

Several facts sharply limit the conclusion. Both models are tiny by current standards. Text quality is poor for both, according to examples supplied by the author. The baseline is a one-layer transformer rather than a suite of modern recurrent and efficient-attention alternatives. Training throughput was not hardware-matched, so the reported T4-versus-CPU rates cannot be treated as an architecture benchmark.

The author also says two important studies remain undone: retention after switching corpora and ablation of the episodic memory bank. Without ablations, it is impossible to know whether the full design is needed. Without scaling curves, there is no evidence that the advantage survives at 10 or 100 times the parameter count.

Practical takeaway

Researchers evaluating PSSA should reproduce the existing comparison before extending it. Pin the code revision, dataset bytes, tokenizer checkpoint, seed, compiler settings, and CPU model. Report wall-clock time, peak memory, loss, accuracy, and generated-token latency separately. Then add modern baselines with comparable parameter counts and optimized implementations.

The next experiment should be an ablation matrix: remove the memory bank, disable plastic updates, change the consolidation rule, and vary recurrent state size. A second useful test would switch domains midway through training and measure both adaptation and forgetting. A third should examine longer contexts, where the theoretical benefit of fixed-size recurrent state is most relevant but information bottlenecks may become severe.

Developers should not deploy the current model merely because its CPU sample is faster. The repository itself says the output is not fluent at this scale. Treat it as a reproducible research object and a source of implementation ideas.

Limitations

All quantitative claims originate with the project author. Public code increases inspectability but is not independent validation. The held-out slice comes from the same broad corpus family, and next-token loss does not measure instruction following, factuality, reasoning, safety, or usefulness.

The comparison also does not establish that transformers have quadratic cost in every production setting; modern systems use caching, optimized kernels, sparse patterns, and other techniques. Conversely, fixed recurrent state can lose detail that full-context attention retains.

The responsible conclusion is narrow: PSSA presents a documented small-scale result worth reproducing. It does not show that a 1.5M-parameter prototype will scale into a general-purpose language model, and the repository is commendably explicit about that boundary.

Share this article

> Want more like this?

Get the best AI insights delivered weekly.

By subscribing, you agree to our Privacy Policy. You can unsubscribe at any time.

> Related Articles

Tags

language modelsstate-space modelsRustopen sourceAI research

> Stay in the loop

Weekly AI tools & insights.