A field expanding toward communication
Brain-to-language decoding now spans text recovery, synthesised voice and expressive avatars. Clinical interfaces are extending the possibilities for people who have lost speech, while non-invasive studies reveal relationships between neural activity and the language people hear, read and imagine. Pretrained models, shared datasets and longer periods of use are bringing these research directions into closer contact.
Our survey connects this work across language tasks, neural recordings, decoding methods, evaluation and practical use. Its organising question is which aspects of language different neural measurements can recover, and how those representations can support expression.
Three conclusions emerge. What a system can recover depends jointly on the participant's task, the neural measurement and the representation it learns. Text, voice and semantic reconstruction develop complementary aspects of expression. And as interfaces improve, feedback, personal adaptation and continued access become increasingly important parts of progress toward communication.

Why the field needs a common frame
A text output can arise from very different scientific questions. One study may identify which passage a participant heard. Another may reconstruct the meaning of a story. A clinical system may recover words the participant attempts to say. Their outputs can look similar on a screen while depending on different neural information, supervision and choices available to the user.
Recording methods add another dimension. They sample different populations of neurons or different consequences of neural activity, with different temporal support and access requirements. A model architecture alone cannot explain which linguistic distinctions are recoverable from those measurements.
These differences make individual results difficult to place without their surrounding conditions. The useful comparison connects a task and recording to a learned representation, then asks what the resulting output enables. This is where a survey can reveal relationships that remain hard to see within any one technical route.
Organising studies around what was actually tested
The survey uses the experimental condition as its unit of comparison: the participant's task, recording, supervision and output, together with the test conditions and use setting. One paper can contain several conditions that support different conclusions. Keeping these choices together helps identify which findings are comparable and which components might transfer to another setting.
The task taxonomy begins with what the participant does:
- Articulated tasks involve actual speech movements, including silent mouthing.
- Inner tasks involve internal production or simulation without an attempt to execute those movements.
- Perceived tasks involve externally presented language, including listening and reading.
Clinical attempted speech remains explicitly described when residual articulation is absent or unconfirmed; it is not automatically treated as Inner speech. The categories organise experimental behaviour; they do not assign language to separate, non-overlapping brain systems.
This task view connects to two further distinctions. Signals describe the information a recording can capture. Methods describe how that information is learned and transformed, separating an intermediate representation from its final presentation as a selected candidate, text, speech or facial/body animation.
Evaluation and practical use complete the map. Quality, Robustness, Efficiency and Credibility organise questions about output fidelity, generalisation, communication costs and the evidence behind a claim. Interaction, support and deployment explain how a capability becomes available to a user. These dimensions are a framework for synthesis, not a new composite score or a universal evaluation standard.
Four patterns across the literature
Different routes preserve different parts of expression
Phoneme and character decoders provide building blocks for word sequences. Acoustic and articulatory models retain structure for synthesising voice. Semantic approaches align neural activity with contextual representations, supporting matching or reconstruction that can preserve aspects of meaning across changes in wording.

These routes make different errors consequential. Word sequences require faithful wording; voice adds intelligibility, prosody and identity; semantic reconstruction must retain relationships and distinctions such as negation. The literature therefore supports several complementary measures of progress. A gain in naturalness or semantic similarity does not by itself establish more faithful transcription.
Reusable learning is valuable when it reduces adaptation
Pretraining contributes at three interfaces: a reusable neural encoder, a speech or language representation used as an alignment target, and a generator that supplies structure for an output. Across neural pretraining, cross-task learning and multi-user models, the common opportunity is to learn useful components from data beyond a single target recording.
The benefit depends on what changes at transfer. New content tests linguistic generalisation; a new participant changes anatomy and recording coverage; a later session changes the mapping within a person. Published comparisons show that personal fitting and reusable models can work together. Performance at a stated amount of target-user data reveals how much prior learning actually saves.
Shared benchmarks reveal progress within a defined problem
Repeated comparisons on common datasets make particular decoding choices directly comparable. The survey brings these results together within their task and evaluation settings. The important agreement is broader than the metric name: studies must also align on the target, test split, permitted inference information and scoring procedure.
This becomes especially important with language-model assistance. A generator can improve a sentence by contributing useful linguistic structure. Controls must establish which distinctions depend on the neural recording. Likewise, recovering a new recording of familiar material and generating text for unseen content test different kinds of generalisation. Clearer benchmarks help locate progress without merging these achievements into one ranking.
Communication improves through a complete interaction loop
Clinical text, streaming speech and longitudinal-use studies extend progress along different dimensions: wording, feedback timing and continued availability. Their shared implication is that an output becomes useful through the actions surrounding it—initiating, inspecting, releasing and repairing a message.

Streaming feedback and sentence-level confirmation arrange these actions differently. Corrections can improve the present exchange and provide accepted targets for later adaptation. Setup and maintenance determine when the interface is available. Accuracy, calibration and interaction costs therefore need to be understood together to assess the benefit to a person communicating.
The gaps lie between established capabilities
The literature establishes that neural recordings can support several forms of linguistic output. It also shows that learned representations can be reused and that communication can continue beyond a short laboratory session. The remaining questions concern how reliably these advances connect across tasks, users and time.
One gap is the relationship between available data and intended expression. Presented language supplies known content and timing; self-chosen messages may have no external transcript or audible reference. Larger source corpora help when they preserve distinctions needed by the target task. The open problem is learning from this supervision while establishing dependable fidelity for the messages a user chooses.
A second gap is predictable transfer. Strong results on one set of recordings do not establish who will benefit elsewhere or how much fitting will be required. Repeated method comparisons, participant-level analysis and independent cohorts provide different pieces of that evidence. Reported benchmarks increasingly specify these conditions, while broad population-level reliability remains less well characterised.
A third gap is sustained usefulness. Faster decoding can still leave substantial correction, preparation and maintenance work. Longitudinal and user-centred assessment must connect technical gains to complete exchanges, availability and participation. These outcomes remain difficult to capture in a single offline score.
Toward richer exchange, with the user in control
The near-term directions follow these gaps: learn representations that remain useful under feasible recording conditions, measure transfer against personal-data budgets, and reduce the work of expressing and repairing a message. Better connections between learning, evaluation and interaction can make existing capabilities more useful for self-chosen communication.
The survey's Beyond framework then asks how the depth of an interface might increase. It proposes five targets: selecting a command, expressing formulated language, conveying intended meaning, sharing an unfolding scenario, and receiving meaningful content through a neural return channel.

The distinctions concern added capabilities. Formulating a sentence differs from conveying an intention independently of its wording. Sharing a scenario adds actors, relationships and change over time. A concept-level return channel would add incoming content that a person can understand and revise without first seeing or hearing it. The illustrations make these differences concrete; they do not demonstrate that the prospective capabilities have been achieved.
This outlook extends the survey's central judgement: richer neural representations matter through the expression they enable. The next stage of progress will also depend on how well a person can recognise, revise and control what is communicated.

