Skip to content

Study compares attention in 13 LLMs, with 17% performing best

Published on October 12, 2025, the study reports experiments training 13 LLMs with different proportions of conventional attention and DeltaNet linear attention. The configuration with 17% attention—two of 12 layers—achieved the best result.

By Wendelmaques ·

Source: Estudo compara níveis de attention em 13 LLMs (x.com). Text prepared with AI from this source.

What happened and what to do

Published on October 12, 2025, the paper reports training 13 LLMs with different proportions of conventional attention and DeltaNet linear attention. In the experiments, 17% attention, placed in two of 12 layers, produced the best observed result. The finding may inform decisions about balancing the two mechanisms in model architectures.

A company that develops or adapts models could apply this approach in a controlled evaluation: compare layer configurations using metrics relevant to its use case, record cost and performance, and track results in a dashboard. This lets architecture decisions draw on tests in the company’s own context without assuming that the study’s best configuration will apply universally.

How the consultancy can help

Wendelmaques can diagnose the opportunity, define an architecture comparison plan, and implement data collection, evaluation, and visualization. The scope can include ongoing operation of the tests on the client’s infrastructure.

Next step

Send a short description of your model, intended use, and team constraints to receive a scoped proposal.

Consulting for your project

Infrastructure review, deployment and ongoing operations, with scope and pricing defined in the proposal.

Quoted per project

Request a proposal

Looking for something else?Frequently asked questionsArticlesContact