Publication
arXiv
Stage
Preprint
What we read
Summary of the paper
Authors
Peter Kulits, Yiqing Xu, R. Kenny Jones, Cordelia Schmid, Jiajun Wu
Universities and research institutions
Not yet supplied in verified metadata; the Brief does not guess.

Institution metadata: OpenAlex record ↗

What they did and found

A 300-prompt BrickBench task with three part limits was used. Eleven new computer programs using artificial intelligence and two basic baselines produced usable Lego brick builds in many cases, but the best human designs remained superior overall part use.

Why it matters

For teams, this shows artificial intelligence can help plan builds and check rules, but people still set the standard for assembly and good part use.

The test used a Lego brick task to see if computer programs using artificial intelligence can follow a prompt and build model that passes test. Many meet basic rules, but human designs still use parts better.

A limit is that the physical test is simple and does not measure strength. Without BrickBench tools, these programs create builds that would not hold up.

What remains uncertain

The physical test uses a basic gravity model; it does not measure real-world strength or long-term durability.

Read the paper PDF ↗

Original sources · 1
  1. BrickBench: Evaluating Agentic Brick Design ↗arXiv · 2026-10-08

Check the original paper for its authors, methods, version and access terms.