Publication
arXiv
Stage
Preprint
What we read
Summary of the abstract
Authors
Zhilin Guo, Boqiao Zhang, Hakan Aktas, Kyle Fogarty, Nursena Koprucu Aslan, Wenzhao Li, Canberk Baykal, Albert Miao, Siyu Hong, Yixiao Liu, Adam Wu, Ashish Kumar Singh, Sakar Khattar, Chenliang Zhou, Weihao Xia, Cristina Nader Vasconcelos, Cengiz Oztireli
Universities and research institutions
Not yet supplied in verified metadata; the Brief does not guess.

What the paper reports

Researchers trained a Telescopic Language Model (TLM) to be valid at every depth prefix, using two forward-backward passes per step and no architectural changes. On a 200M proxy suite, a single TLM run yielded a valid model across all prefixes and reduced the quality-budget area by about 43-44% versus fixed-exit approaches, while matching full-capacity performance with ~12% lower GPU cost per run.

Why it matters

The study suggests training objectives, not nesting alone, enable models to operate across varying budgets while preserving quality, offering a path to more flexible deployment without re-architecting models.

Telescopic Language Models (TLMs) are designed to function well at multiple depths within the same model, by training with random truncations of the capacity axis and the full-capacity target. This approach allows one model to serve a spectrum of compute budgets without separate training for each point.

In tests on a 200M proxy data setup, a single TLM run produced a usable language model at every depth prefix and showed lower total training cost per run compared to fixed-exit alternatives, though performance could vary by depth and task.

What this does not tell us

This is an abstract-only, preprint study. Claims are based on a specific 200M proxy suite with abstract-only evidence; results may not generalize to all data scales or real-world deployments.

Original sources · 1
  1. Telescopic Language Models ↗arXiv · 2026-09-28

Check the original paper for its authors, methods, version and access terms.