September 2026

Receiver-Conditioned Latent Communication gives 94% CacheBack

  • Maximillian Rossi
  • Prajwal Raghunath
  • Haoqing Xuan
  • Yusen Zhang
  • Eugene Wu

DAP LabColumbia University

TLDR; Agents can share internal model state instead of generating text messages. Sending all of it, however, overwhelms the receiver.

The receiver asks for what it needs. CacheBack uses attention to that request to select which state gets sent. Simple, robust, and training-free.

At 16× compression, CacheBack passes on just 6.25% of sender positions.

Selected source positions can be sent as token IDs; latent steps must be sent as continuous vectors. The receiver prefills both to build its own state.

CacheBack on a coding task

36 seconds · 4K · Recorded coding task

Fixing a Django bug with seven coding agents

Seven Qwen3-8B workers inspect the code and send information to a coordinator, which writes the Django fix. The video compares selected latent state with generated text messages. CacheBack produces the patch 4.41× faster than with text communication on this case. Both runs make the same one-line code change and pass all 88 tests.

One recorded case: CacheBack 25.66 s; text 113.21 s. Separate recorded runs are aligned at their start. Startup and test grading excluded. Silent video; playback speed varies. Open in full window

Results

FanOutQA

Questions require combining evidence across Wikipedia pages. We split at least 120K tokens among three parallel senders; a receiver answers from their messages. Strict accuracy requires every reference-answer group. Higher is more accurate; left is faster. Each panel fixes the receiver model. Text labels give sender size; CacheBack uses same-size senders. 4× compression retains ¼ of sender positions; 16× retains ¹⁄₁₆.

● ━ CacheBack · labels show compression factor▲ ┄ Text · labels show sender size× Question only
FanOutQA, Qwen 3 8B: CacheBack 4× compression reaches 55.3% strict accuracy at 192 seconds, compared with 40.7% at 612 seconds for same-size text; weaker settings are also plotted.
Qwen: 4× compression improves accuracy and completion time over 8B text; 1.7B text is slightly faster but less accurate.
FanOutQA, Nemotron Nano 2 12B: CacheBack 8× compression reaches 50% accuracy near 102 seconds versus 38.7% near 131 seconds for same-size text. A smaller text sender is faster at lower accuracy.
Nemotron: 8× compression improves on 12B text; 4B text is faster but less accurate.

LongBench v2 Easy

Long-document questions test whether evidence survives repeated handoffs. We split each 100K-246K-token document into four equal parts. Each sender reads its part plus the previous message; a final receiver answers from the fourth message. Compression factors retain a fraction of accumulated context (4× keeps ¼); fixed budgets (32K, 64K) cap message positions. Axes and text-size labels follow FanOutQA.

● ━ CacheBack · labels show compression factor■ ━ CacheBack · labels show fixed position budget▲ ┄ Text · labels show sender size× Question only
LongBench v2 Easy, Qwen 3 8B: both relative and fixed CacheBack budgets include settings faster and more accurate than same-size text. Stronger compression eventually lowers accuracy.
Qwen: 4× compression gives the highest plotted accuracy; stronger compression loses accuracy.
LongBench v2 Easy, Nemotron Nano 2 12B: CacheBack 8× compression reaches 48% accuracy near 133 seconds, while fixed-budget settings show a different tradeoff; smaller text senders can be faster.
Nemotron: relative 8× compression beats 12B text on both axes; 4B text remains faster but less accurate.

Lines join the best accuracy-time tradeoffs within each channel; faded points are dominated settings. × means the receiver gets only the question. 50 concurrent tasks on 8 H100s; accuracy averaged over three receiver draws. Completion includes queueing. Full setup and results in the paper.

Interactive demo · Unstructured documents

Multi-hop document Q&A

Open full window

In multi-agent question answering, a large document collection can be split across agents, each with its own context. Answering a question requires combining information across those contexts.

With text communication, agents write messages summarising what they read. CacheBack lets them pass latent thoughts together with a subset of internal state selected for what the next agent needs. This example shows selected-state handoffs: three agents read separate documents in sequence, and a fourth agent produces the final answer.

The top row shows each agent’s input document: a booking confirmation, a location-change notice, then an ID policy. Agents 2 and 3 also receive the previous agent’s handoff; Agent 4 answers from the final handoff alone. The bottom row shows the messages generated by the text agents.

Question

Where should Maya collect her pass on Friday, and what ID should she bring?

Press Run to watch the handoffs.

Use CacheBack

Install rclc, bind your agents, and pass the receiver’s request.

Code and documentation
Python
import rclc

sender = rclc.bind(
    model, tokenizer,
    messages=sender_history, backend="hf",
)
receiver = rclc.bind(
    model, tokenizer,
    messages=receiver_history, backend="hf",
)
await rclc.transfer(
    sender, receiver,
    "Who owns the Cedar booking?",
)
inputs = receiver.pop()

Model, tokenizer, and histories loaded; run in an async function or notebook. Runnable quickstart

THE PAPER

Receiver-Conditioned Latent Communication gives 94% CacheBack

Maximillian Rossi, Prajwal Raghunath, Haoqing Xuan, Yusen Zhang, and Eugene Wu.

Cite this work

@misc{rossi2026cacheback,
  title = {Receiver-Conditioned Latent Communication
           gives 94\% CacheBack},
  author = {Rossi, Maximillian and Raghunath, Prajwal
            and Xuan, Haoqing and Zhang, Yusen
            and Wu, Eugene},
  year = {2026},
  eprint = {2609.32046},
  archivePrefix = {arXiv},
  primaryClass = {cs.AI},
  url = {https://arxiv.org/abs/2609.32046}
}