$npx skillfedfor your agent
RESEARCH

Splitting long-range windows across KV heads beats duplicating them at lower cost

on: CoWindow Attention: Full Causal Coverage Is a Collective Property

Every attention head in a standard transformer sees the full causal history. That redundancy is the point — or so the assumption goes. CoWindow Attention (CoWA) challenges it directly: full causal coverage does not require every head to duplicate the same long-range access. It only requires that, across all heads together, every prior token is visible to at least one.

The mechanism is simple by design. All KV heads share a near-diagonal window covering recent tokens and a prefix-sink window anchoring the sequence start. The remaining causal history is partitioned into non-overlapping long-range intervals, one assigned to each KV head. No router, no learned indexer, no content-dependent selection — just position arithmetic applied identically during training and inference.

The ablation that makes the argument concrete trains models with one, two, four, and eight distinct long-range windows distributed across eight KV heads, holding per-head window widths fixed. Associative recall at 8K climbs monotonically: 21.32%, 32.92%, 52.32%, 89.73% as duplication gives way to complementarity. Full collective coverage nearly matches FullAttn's 89.97%. Duplicated windows at matched widths are not a substitute for distributed ones.

At scale, the efficiency case is substantial. During 32K long-context training at 14B parameters, CoWA cuts total training FLOPs by 28.5% relative to full attention. The savings are modest at 4K context (3.1%) and grow with sequence length, which is exactly where you want them. On an 8-GPU H100 operator benchmark at 128K tokens, CoWA also reduces decoding latency and uses less peak memory per rank during decoding than full attention — and unlike MoBA and DSA, it carries no routing or indexing overhead because its sparse pattern is derived directly from position and global KV-head index.

Model quality holds. At 14B, CoWA scores 72.70 on knowledge benchmarks versus 72.32 for full attention, and 64.87 versus 64.46 on reasoning. At 32B via continued training, the margins are similarly small and go in both directions depending on the task. Long-context retrieval on RULER at native 32K stays within 0.3 points of full attention at both scales. The 128K extrapolation results are noisier, with CoWA ahead at 14B and marginally behind at 32B.

The design aligns naturally with tensor parallelism: global KV-head indexing ensures that complementary long-range windows are preserved across ranks rather than scrambled by sharding. This is not an afterthought — it is part of why the pattern stays regular and cheap to execute.

What CoWA does not claim: identical head-level behavior to full attention. Each head applies softmax over its own visible set, so the per-head outputs differ even when aggregate coverage is the same. The paper is careful about this. Collective coverage is a property of the layer, not of any individual head, and the model learns to use that structure end to end.

Distributing long-range attention across KV heads rather than duplicating it cuts 32K training FLOPs by 28.5% at 14B with no meaningful quality loss.

Sources & links