T²Mem:

Learning Test-Time Memory for Robotics

Yize Liu1 · Huang Huang1 · Yining Hong1 · Zijian Du2 · Zhi Cao3 · Li Fei-Fei1 · Jiajun Wu11 Stanford University   2 NVIDIA   3 University of Michigan, Ann Arbor

yizeliu@stanford.edu

Memory as an intrinsic capability: a robot policy learns what to remember and how to use it, without auxiliary memory models or memory-specific annotations.

VideoUnmask

Recall a hidden object.

PickXtimes

Count repeated actions.

PatternLock

Recall a demonstrated sequence.

Task demonstrations from RoboMME. Red border: demonstration phase.

56.83%

Mean success on RoboMME

+38.90pp

Over the memory-free π0.5 baseline

16tasks

Across four memory categories

Abstract

Memory-dependent robotic manipulation requires policies to use information that is no longer available in the current observation. Retaining history alone is insufficient: memory must preserve information that supports future actions. One challenge is whether a memory-free foundation model can learn to retain and use historical information from action demonstrations alone, without external memory support. We introduce T²Mem, a framework that develops this capability within a pretrained vision-language-action policy, without external reasoning models or memory-specific annotations. T²Mem uses test-time training to encode observation history into compact fast weights through online self-supervised updates, avoiding repeated processing of the full history. An observation-grounded interface extracts vision-language information for memory formation and supplies retrieved context to the action expert. Action supervision shapes what the memory learns to retain and use, while alternating memory-policy learning gives each component a fixed counterpart during optimization. Across 16 RoboMME tasks, T²Mem improves average success from 17.93% to 56.83% over the memory-free base policy and outperforms the recurrent-memory methods reported in the benchmark, while controlled profiling indicates at least 3x inference speedup over explicit methods.

01 / THE METHOD

Observation-Grounded Memory

Observe → read memory → act → update memory.

Observation → interface
Observation + instruction↓Vision-language backbone
Current scene features
OBSERVATION-GROUNDEDMemory interfaceObservation features
read ←write →
EPISODE STATEFast weights W
SELF-SUPERVISED WRITEL = mean[(fW(K) − V)²]
↓Action expert↓History-conditioned action

Illustrative features · Blue: observation · Violet: retrieved history · Orange: memory update

01

Grounded in observations

Memory writes come from visual-language observations.

02

Adaptive self-supervision

Write strength adapts to new observations.

03

Independent episode state

Fast weights update online and reset each episode.

Explore the full architecture figureClick to expand the paper diagram
Detailed T²Mem architecture: VLM, memory interface, fast-weight read and write paths, and action expert.
Architecture of T²Mem. (a) Memory connects the vision-language model (VLM) and action expert alongside the direct vision-language pathway. (b) The observation-grounded interface extracts vision-language features, fuses memory readouts, and supplies attention context to the action expert. (c) Self-supervised updates encode history in fast weights.

Architecture of T²Mem

As shown in the architecture figure(a), T²Mem builds on π0.5 and comprises a VLM backbone, a memory module, and an AE. The memory sits between the VLM and AE. Its memory interface extracts vision-language features and combines them with retrieved history to condition the AE. This pathway complements the original VLM-to-AE connection, which preserves direct access to the current scene.

We distinguish slow parameters, learned across episodes, from fast state, updated within an episode. Slow parameters include the policy, interface, memory projections, fusion gates, and fast-weight initialization W0. They learn how to encode and use history, which responsible for "how to memorize" Fast state Wt consists of the current weights and biases of the memory networks and carries episode-specific information, which is "what to memorize". Tasks share slow parameters, while each episode starts from W0 with an independent fast state. During inference, all slow parameters remain fixed, while fast parameters undergo continuous self-supervised updates. During training, all parameters are unfrozen except those of the vision and language encoders.

Observation-Grounded Memory Interface

The memory interface uses learned query tokens to aggregate VLM features into a compact state summary for memory writes, retrieval, and action conditioning (the architecture figure(b)). This leverages pretrained vision-language representations without updating memory over the full visual token sequence. At depth l, interface representations Et(l) extract information from vision-language features Zt(l):

Ut(l)=Attn(Qif(l)(Et(l)),KVL(l)(Zt(l)),VVL(l)(Zt(l))).

An attention mask restricts interface queries to valid vision-language keys, excluding action, proprioceptive, and other interface tokens while preserving the AE's original connections. This grounds memory in observations and blocks a direct action-history shortcut that could reduce imitation loss through action extrapolation rather than task-state tracking.

The summary Ut supports both retrieval and writing (the architecture figure(c)). Omitting layer/head indices, normalization, and positional encoding, retrieval uses an observation-conditioned query:

Qt=Utθq,Rt=fWt(Qt),

where fWt is the fast-weight memory and Rt its history-conditioned readout. For writing, dedicated projections form key–value associations from the same summary:

Kt=Utθk,Vt=Utθv.

At scheduled write steps, a self-supervised update encodes these associations into the fast weights:

ℒmem,t&=mean[(fWt(Kt)−Vt)2],Wt+1&=Wt−ηteff∇Wtℒmem,t,

where ηteff=ηtβt includes adaptive scaling (Adaptive Self-Supervised Memory). Reads precede writes, so new observations affect subsequent decisions. Memory projections θq, θk, and θv are separate from the interface attention projections and learned through the outer action objective; inner updates modify only episode-specific fast weights.

A channel-wise gate fuses retrieved history with the current vision-language summary:

U~t(l)=Ut(l)+tanh(α(l))⊙Rt(l).

Residual and feed-forward transformations then produce Ct(l). Following RoboTTT , the gate initially limits memory's contribution to protect pretrained visuomotor capabilities; action supervision learns its channel-wise fusion strengths.

The enriched interface supplies keys and values to subsequent action-token attention, alongside current vision-language features, proprioception, and action context. An interface update is consumed after its layer, not by that layer's completed attention computation. The AE thus combines VLM perception and retrieved history to predict actions.

Inside Test-Time Training

Read with a forward pass. Write with a self-supervised gradient.

READ · Retrieve memoryProject the interface → Qₜ
Interface projections, fast-weight forward pass, and self-supervised backward update Interface UₜVL features acquired Fast weights · fWₜForward pass · Wₜ unchanged θqQₜRₜMemory context θkKₜθvVₜ · targetComparefWₜ(Kₜ) − VₜPrediction error
Rₜ = fWₜ(Uₜ θQ)Read: retrieve history without changing W.

Schematic network and feature colors. Read → Write → Repeat · Only fast weights update during inference.

Test-Time Training

TTT updates a model at inference using a self-supervised objective constructed from its inputs. In the fast-weight formulation, a neural model fWt encodes episode history in its parameters Wt . Given an input representation xt, slow parameters θ produce queries, keys, and values (qt,kt,vt). A standard associative objective is

ℒmem(W;xt)=12‖fW(kt)−vt‖22,

with memory readout and update

mt&=fWt(qt),Wt+1&=Wt−ηt∇Wℒmem(W;xt)|W=Wt.

The action policy is conditioned on the readout, πθ(at∣ot,ℓ,mt). We adopt a read-before-write convention with inner-loop step size ηt; setting ηt=0 skips a write. The slow parameters θ and initialization W0 are learned through the outer action objective. At deployment, θ remains fixed, while fast weights update without expert action labels and reset to W0 at each episode boundary.

Adaptive Self-Supervised Memory

Fast memory encodes history in an online-updated nonlinear mapping (the architecture figure(c)). The observation-grounded interface provides queries, keys, and values for retrieval and self-supervised association learning. This update requires no memory-content labels. The initialization, projections, and step-size parameters are learned through outer action supervision and fixed at deployment.

To limit repeated reinforcement from correlated observations, we adapt write strength using the alignment between the descent direction and Wt−W0, together with the normalized reconstruction residual. These signals determine a scale βt∈[0.1,1] that attenuates aligned updates while preserving stronger updates for less aligned, poorly reconstructed inputs:

Wt+1=Wt−wtηtβt∇WLmem,t(W)|W=Wt,

where wt is the binary write mask and ηt is a curvature-calibrated step size with a learned positive multiplier. The non-adaptive control fixes βt=1 without changing the step-size rule.

Reads precede writes, so each update affects only subsequent predictions. The fixed-size fast state avoids a growing history buffer. Architecture, scaling, and gradient details are provided in Adaptive Memory in the paper appendix; training and inference procedures appear in Training Details in the paper appendix.

02 / LEARNING

Alternating Memory–Policy Learning

Learn what to remember. Then learn how to use it.

Stage 1
Action supervision updates the policy
MemorySlow parameters
PolicySlow parameters
Action lossExpert demonstrationsSupervision
Memory inactive

Blue: slow-parameter learning · Orange: episode-local fast-weight updates

↻   Alternate 2A ↔ 2B. Fast weights update in both.

Self-supervised association learning does not by itself ensure decision-relevant memory, and the policy must learn to use the new historical representations. In joint training, memory updates change the context presented to the policy, while policy updates change the action gradients that guide memory learning. The two modules may therefore continually adapt to each other's changing representations, making a stable memory-to-action mapping harder to learn. We address this coupling by alternating their updates, holding one slow-parameter group fixed while optimizing the other. The animation above outlines the training stages, which is, learning to remember and learning to act are separate.

Initialization and memory-free adaptation. We initialize from pretrained π0.5 (Stage 0) and adapt the policy without memory (Stage 1). This establishes task-specific visuomotor skills before introducing the memory pathway.

Stage 2A: memory learning. We fix the VLM/AE and train the memory mechanism to extract, store, and retrieve history useful for action prediction. Expert action supervision passes through the frozen AE to shape both current retrieval and earlier memory writes. The fixed policy provides a stable decision-making counterpart, encouraging memory representations that serve its control needs.

Stage 2B: memory-conditioned policy learning. We fix the memory mechanism and adapt the policy to combine current observations with retrieved history. This phase adjusts both the vision-language features supplied to memory and their use by the AE. Only memory slow parameters are frozen: episode-specific fast states continue to accumulate observations through online updates.

Alternating schedule. We repeat these phases so that memory adapts to the current policy, the policy learns to use the resulting representations, and subsequent memory learning responds to the updated policy. This repeated adaptation differs from training memory once and then fitting a policy to it. Both phases use expert action supervision without memory-content or task-progress labels. The action objective, parameter groups, sequence supervision, and training schedule are detailed in Training Details in the paper appendix.

03 / RESULTS

Results on RoboMME

Table 1 from the paper: full success rates on all 16 RoboMME tasks, including human performance, symbolic, perceptual and recurrent memory baselines, and T²Mem. T²Mem achieves 56.83% average success.
The standard evaluation uses 50 test episodes per task and three evaluation seeds. Baseline scores are quoted directly from the RoboMME paper; human and privileged Oracle results are reference points rather than deployable competitors. T²Mem improves average success from 17.93% to 56.83% over the memory-free base policy. Click the table to enlarge.

Performance on Memory-Dependent Tasks

We first examine whether T²Mem enables effective control on memory-dependent tasks. The table above summarizes the RoboMME evaluation (including no memory baseline). Among methods without privileged information, T²Mem ranks within the top three on most reported tasks, with the highest success rates on VideoRepick, InsertPeg, and RouteStick. It also substantially improves over the memory-free π0.5 and the evaluated recurrent-memory baselines.

Performance is strongest on tasks that primarily require direct retrieval of an earlier cue or tracking repeated events, such as VideoUnmask and StopCube. In contrast, success remains lower on VideoUnmaskSwap and ButtonUnmaskSwap, where the policy must track swaps and update object–location associations rather than simply recall a stored binding. This contrast suggests a distinction between retaining information and reasoning over it. As a single-model approach, T²Mem relies on the underlying π0.5 policy for visual feature extraction and temporal reasoning, without an external reasoning model. Its memory supplies historical evidence but does not, by itself, confer the ability to infer how that evidence changes through subsequent events. The weaker swap performance may therefore reflect limitations in the base policy(π0.5)'s temporal reasoning, which improved memory retention alone cannot resolve.

Low-level control imposes a separate limitation on tasks such as InsertPeg. Our qualitative observations on InsertPeg indicate that the policy can identify and approach the intended target yet fail during the final insertion. Thus, terminal success reflects both memory-dependent decision-making and execution precision. The strong relative improvement on InsertPeg, despite its modest absolute success rate, highlights the importance of distinguishing these failure sources.

Ablation Studies

Paper ablation figure: memory architecture, alternating learning, and adaptive memory writing.
Ablation studies of T²Mem. (a) Memory architecture and (b) training schedule: success rates on VideoUnmask and MoveCube, with a memory-free π₀.₅ reference. Mean denotes the average across the two tasks. (c) Adaptive writing: reconstruction NMSE of early-memory K/V bindings over subsequent observations, averaged over 20 VideoUnmask Hard episodes. Lower is better; β = 1 disables adaptive scaling but retains step-size calibration.

Observation-grounded memory architecture. The comparison places memory either in our observation-grounded interface or within the action expert, where reads and writes derive from action-expert representations. T²Mem achieves 84% and 71% success on VideoUnmask and MoveCube, compared with 30% and 26% for the action-expert-side implementation. This supports using semantic vision-language features to form memory for these tasks.

Alternating memory and policy learning. The ablation models are trained independently on each task with identical data, batch size, and initialization. Under the same training budget, jointly updating memory and policy performs similarly to the memory-free baseline. Alternation holds one component fixed while the other adapts, giving memory learning and memory-conditioned control a stable counterpart.

Preserving earlier associations. Starting from shared visible-demonstration memory, the adaptive-writing experiment replays 300 subsequent real frames across 20 long VideoUnmask Hard episodes. Architecture, write cadence, and step-size normalization remain unchanged. Adaptive writing lowers reconstruction error on the initial K/V bindings by 18.8%–46.9%, supporting reduced interference with earlier memory.

04 / INSIDE THE MEMORY

Does the Policy Use Its Memory?

Same scene. Different histories. Different choices.

Same observation. Different histories.

VideoUnmask · counterfactual expert demonstrations

Front viewWrist view
History ABlue cube in the first location.
Front viewWrist view
History BBlue and green positions exchanged.

Same instruction: “Watch the video carefully, then pick up the container hiding the blue cube.” The container layout stays fixed; the earlier color locations change, so the robot must grasp a different container. Both clips include the full grasp.

Expert demonstrations illustrating the counterfactual task, not T²Mem policy rollouts.

CHANGE ONLY THE MEMORY
1 · DEMONSTRATIONSame episode’s demonstration
2 · MEMORYMatching history
3 · CURRENT VIEWSame hidden scene

True target: A

4 · CHOICE (ILLUSTRATION)Matching history guides selection of A.
ATrue target
BOther target
COther target
✓ Correct target
How this experiment works

Correct: use memory built from this episode’s own demonstration.

Schematic illustration: A / B / C denote possible targets, not labels supplied to the policy. Choices are schematic, not recorded trajectories; random choices do not represent measured policy probabilities.

50 counterfactual pairs × both directions × 3 seeds. Instruction, environment goal and diffusion seed stay fixed.

VIDEOUNMASK / MEMORY CONTENT
100% success

300 / 300 successful trials

The right history guides the action.

Paper figure comparing online writes, correct versus conflicting or empty memory, and write timing.
Online memory supports history-dependent decisions. (a) Disabling fast-weight writes reduces success across three tasks. (b) Correct, conflicting, and empty memories produce markedly different outcomes on VideoUnmask. (c) Pre-occlusion writes match full-demonstration performance with half the updates, whereas post-occlusion writes perform substantially worse. Success rates are pooled over three seeds under each panel’s evaluation protocol.

Online updates are necessary. Disabling episode-local fast-weight writes retains memory reads and the learned initialization, yet reduces success by 82.67 percentage points on VideoUnmask, 64.67 on SwingXtimes, and 41.33 on MoveCube. The fixed policy and initial memory alone cannot sustain normal performance.

The policy uses the content of history. For 50 counterfactual VideoUnmask pairs, the instruction, environment goal, and diffusion seed remain fixed while memory is replaced. Across both pairing directions and three seeds, correct, conflicting, and empty memories yield 300/300, 1/300, and 101/300 successes. Conflicting history is more damaging than empty memory, supporting content-specific use of the stored information.

Useful writes occur while the cue is visible. With equal budgets of eight writes, pre-occlusion memory achieves 76.67% success, versus 20.67% after occlusion. Eight pre-occlusion writes also match all 16 demonstration writes. This localizes decision-relevant evidence to the interval when the target remains visible.

Inference Efficiency

Paper figure showing memory-stage latency and total inference latency for T²Mem, the highest-scoring RoboMME baseline, and MemER.
Controlled-workload inference efficiency. Left: latency introduced by memory only. Right: end-to-end latency for an inference. T²Mem, FrameSamp+Modul, and MemER are profiled on RTX A5000 GPUs at batch size one. Timing uses device synchronization after warm-up, excluding initialization, compilation, communication, and environment execution. T²Mem is approximately 3.0× faster than FrameSamp+Modul and 69.2× faster than MemER under this profiling setup.

Compact history enables faster decisions. T²Mem stores and retrieves history in latent fast weights within a single model. It avoids repeatedly processing the full history and does not invoke an external autoregressive reasoning model. Under the controlled workload, foreground computation is approximately 3.0× faster than FrameSamp+Modul and 69.2× faster than MemER.

Interpreting the timing. The left panel isolates memory-stage computation; the right includes the full profiled inference computation. Measurements use synchronized, warmed-up execution on RTX A5000 GPUs at batch size one. They exclude initialization, compilation, communication, and environment execution, so the comparison measures model computation rather than the duration of a complete robot episode.

Citation

@misc{liu2026t2memlearningtesttimememory,
  title={T$^2$Mem: Learning Test-Time Memory for Robotics},
  author={Yize Liu and Huang Huang and Yining Hong and Zijian Du and Zhi Cao and Li Fei-Fei and Jiajun Wu},
  year={2026},
  eprint={2609.36720},
  archivePrefix={arXiv},
  primaryClass={cs.RO},
  url={https://arxiv.org/abs/2609.36720},
}