WEDNESDAY · SEPTEMBER 30, 2026 · ISSUE 21 · YESTERDAY No. 21
Daily News — 2026-09-30
781 papers indexed on arXiv·~28,000 packages released on PyPI·50 papers surfaced by Hugging Face
None of it is in your agent's weights.
Spotlight
Composing an explicit, editable score before rendering audio demonstrably improves what expert listeners hear—and makes the composition itself available for revision.
Writing a score before rendering audio turns out to matter. YuE2's central claim is that making melody, harmony, rhythm, and form explicit in a readable intermediate representation—before any acoustic generation…
A C++-patched Firefox that builds synthetic identities from the engine out, not the JavaScript layer up—the right level to fight detection.
The central claim here is architectural, not cosmetic: web agents fail at the browser layer, not the model layer. Most automation frameworks bolt anti-detection measures onto a standard browser after the fact—JavaScript…
Papers
Composing an explicit, editable score before rendering audio demonstrably improves what expert listeners hear—and makes the composition itself available for revision.
YuE2 writes a readable score before generating audio, beating Suno v4.5 in expert listening and nearly matching Suno v5 on a best-of-8 basis.
Instruction-following feedback dominates multi-teacher distillation gradients; rescaling by batch-wise spread recovers the mathematics gain that label routing loses.
DN-MOPD rescales each specialist's feedback by its measured spread, fixing the imbalance that lets instruction-following dominate multi-teacher distillation.
Wrapping a VLA inside an executable world-and-policy program beats the same VLA alone by attributable, reproducible margins—the mechanism is the key result.
HexaAnything wraps VLA/WAM policies in code-represented task state so verified execution traces feed back as training data, beating direct VLA on unseen tasks.
Four deliberately non-aggregable scores expose failure modes that a single response rate would bury in any multi-party spoken assistant evaluation.
Duplex-MPE is a benchmark of 2,000 scenarios testing whether a full-duplex speech assistant answers, stays silent, or stops in multi-party conversations.
Distributing long-range attention across KV heads rather than duplicating it cuts 32K training FLOPs by 28.5% at 14B with no meaningful quality loss.
CoWA splits causal-history access across KV heads for full collective coverage, cutting 128K training latency 7-8x while tracking FullAttn accuracy at scale.
Encoder-free multimodal models are predicted to match encoder-based ones around 10²³ FLOPs—and the decoder's internal reorganization explains why.
Scaling-laws analysis predicts encoder-free MLLMs close the multimodal gap with encoder-based models at ~10^22 FLOPs, within practical pretraining budgets.
9 more tool picks in this edition
Every pick in the Wire gets the same treatment: read, verified, and given a written verdict.
TaH2 turns the looped transformer's steeper scaling slope into actual accuracy gains by learning per-token iteration depth jointly with the backbone, not as an afterthought.
TaH2 adds a learned iteration decider to looped transformers, lifting the accuracy-compute slope 53% over a non-looped baseline on AIME benchmarks.
A playbook distilled from a strong agent's failures can make a cheap agent outperform the strong agent running blind — and the numbers hold on a physical robot.
A method that distills a strong agent's robot manipulation experience into a reusable playbook, improving success from 37.3% to 64.0% in real-world tasks.
Thinking longer only helps when the answer is already reachable; this paper proves the distinction empirically and trains small models to act on it.
FlyBy trains 4B/8B models to detect knowledge bottlenecks mid-reasoning and query stronger models, letting FlyBy-4B beat Qwen3-14B at 2.7x lower serving cost.
Tools & packages
A C++-patched Firefox that builds synthetic identities from the engine out, not the JavaScript layer up—the right level to fight detection.
An open-source AI agent with its own browser, designed to avoid being blocked on the web.
A dense, VRAM-grounded catalogue of 27 uncensored security LLMs that separates domain-fine-tuned from abliterated models and tells you exactly what hardware each one needs.
Curated list of uncensored and cybersecurity-fine-tuned models for offensive security and red-teaming tasks.
Technically interesting MCP-to-SolidWorks bridge, but the installation scripts pipe from unverified third-party domains - do not run them.
MCP server that connects an AI assistant to a live SolidWorks instance to sketch, extrude, fillet, export STEP/STL, and generate macros.
Genuine UI depth for ChatGPT plugins, but the host-specific coupling is the whole story — this is not portable MCP.
MCP extensions for ChatGPT that make plugins feel like first-class native features rather than third-party add-ons.
A Claude Code plugin that turns a one-line prompt into a working game mod by encoding twelve engine playbooks, a full asset pipeline, and hard safety refusals into a single agentic loop.
A toolkit of skills, tools, and MCP integrations that lets Claude Code recon, reverse-engineer, and mod PC games with generated art, 3D, and audio.
A Claude Code skill that turns a video and transcript into timeline-ready motion-graphic B-roll using pure-function spring animation and headless Chromium capture.
Motion-graphics skills repo for Claude Code and Codex - covers animations, transitions, and visual effects.
9 more paper picks in this edition
Every pick in the Wire gets the same treatment: read, verified, and given a written verdict.
A structured Claude Design skill that animates UX rationale frame-by-frame — closer to a typed contract than a prompt.
An animated blueprint drawing that walks through a before-and-after UX redesign step by step.
A credible Ascend port of DeepGEMM that hits 99.8% hardware utilization on BF16 and covers the full MoE kernel stack DeepSeek models actually need.
DeepSeek releases a clean, efficient GEMM kernel library targeting Huawei Ascend NPUs.
A complete self-hosted Claude client for 2007 Nokia hardware, built around the real TLS incompatibility that makes every other approach fail.
Because a 2007 Nokia can't Google anymore, this unofficial J2ME app + Go server gives it Claude instead.