Publication
arXiv
Stage
Preprint
What we read
Summary of the paper
Authors
Sophie L. Wang, Amil Dravid, Rulin Shao, Kevin Farhat, Sewon Min, Alexei A. Efros
Universities and research institutions
Not yet supplied in verified metadata; the Brief does not guess.

What they did and found

What happened: A study found that just fixing the first few words a base AI model sees can move its math and coding scores closer to those of rewards-trained models. The result depends on the model, task, and data cues; safety behaviors can shift too.

Why it matters

For teams, this suggests careful data curation may reduce extra training. But benefits are uneven and depend on context, so test locally and expect variation.

One finding: fixing the first two words can change a base model’s answers, sometimes making math and coding results closer to those from reward-based training.

The lesson is to test cues in your own setup, since results depend on model type and problem, and some safety responses may change.

What remains uncertain

Effects differ by model and task; not all base models respond to cues, and gains may not transfer to real deployments.

Read the paper PDF ↗

Original sources · 1
  1. Base Models Can Reason By Taking a Cue From Training Data ↗arXiv · 2026-10-05

Check the original paper for its authors, methods, version and access terms.