One Transformer block reused to match full depth with few expert parts
A smaller model uses a depth-guided bank of expert parts to imitate deep behavior.
- Publication
- arXiv
- Stage
- Preprint
- What we read
- Summary of the paper
- Authors
- Adrian Bulat, Yassine Ouali, Georgios Tzimiropoulos
- Universities and research institutions
- Not yet supplied in verified metadata; the Brief does not guess.
Institution metadata: OpenAlex record ↗
What they did and found
Problem: deep vision models need many layers to capture features. Approach: reuse a single Transformer block many times and select from a small pool of expert feed-forward networks at each depth, based on position. Finding: similar accuracy with similar compute.
Why it matters
This could let teams build smaller models that still act like they have extra depth, but results depend on data and training setup.
A single module runs repeatedly; a small set of expert parts adapts output at each depth. The method keeps computation similar to a basic block while adding depth-like changes.
In practice, you pick a deployment depth and reuse the same block, letting the depth decide which tiny feed-forward network from the expert bank to use, instead of adding layers.
What remains uncertain
Results depend on model type and training; using flexible depth and distillation can change outcomes in different scenarios.
Original sources · 1
- One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed Experts ↗arXiv · 2026-10-08
Check the original paper for its authors, methods, version and access terms.
What is your take?
Ask a question, add useful context or share a different perspective. Keep the conversation respectful and grounded.