Writing a score before rendering audio turns out to matter. YuE2's central claim is that making melody, harmony, rhythm, and form explicit in a readable intermediate representation—before any acoustic generation…
The central claim here is architectural, not cosmetic: web agents fail at the browser layer, not the model layer. Most automation frameworks bolt anti-detection measures onto a standard browser after the fact—JavaScript…
DN-MOPD rescales each specialist's feedback by its measured spread, fixing the imbalance that lets instruction-following dominate multi-teacher distillation.
Read by the desk
↑140hf upvotes
RESEARCHSelf-Evolving Coding Agents: From Digital Programs to Physical-World Intelligence · arXiv · Sep 30
HexaAnything wraps VLA/WAM policies in code-represented task state so verified execution traces feed back as training data, beating direct VLA on unseen tasks.
Duplex-MPE is a benchmark of 2,000 scenarios testing whether a full-duplex speech assistant answers, stays silent, or stops in multi-party conversations.
Read by the desk
↑84hf upvotes
RESEARCHCoWindow Attention: Full Causal Coverage Is a Collective Property · arXiv · Sep 30
CoWA splits causal-history access across KV heads for full collective coverage, cutting 128K training latency 7-8x while tracking FullAttn accuracy at scale.
Read by the desk
↑60hf upvotes
RESEARCHHow Far Are We from Removing the Visual Encoder? Scaling Laws for Encoder-Free Multimodal Pretraining · arXiv · Sep 30
Scaling-laws analysis predicts encoder-free MLLMs close the multimodal gap with encoder-based models at ~10^22 FLOPs, within practical pretraining budgets.
A method that distills a strong agent's robot manipulation experience into a reusable playbook, improving success from 37.3% to 64.0% in real-world tasks.
Read by the desk
↑35hf upvotes
RESEARCHKnowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge · arXiv · Sep 30
Disaggregated quantization uses separate formats for prefill and decode, yielding 32+ point accuracy gains and 1.78x faster first-token latency on 27B models.
InternW0-Δ unifies visual dynamics, 4D geometry, and action generation in a Mixture-of-Transformers WAM trained on 20K+ hours of open robot and egocentric data.
WROP pairs a 1.5M-sample object-permanence corpus with a 300-question blind Elo exam; their 16B PWM-WROP ranks first among continuation video world models.
Read by the desk
↑193hf upvotes
RESEARCHAgent-Editing World Model: Rethinking World Modeling for LLM Agents · arXiv · Sep 26
AEWM edits noisy reasoning and actions in task history rather than predicting tool responses, improving agent scores by 3.2–6.7 points across six benchmarks.
Read by the desk
↑16hf upvotes
RESEARCHParts-of-Speech as Emergent Categories in SAE Latent Space · arXiv · Sep 26
A study finding that part-of-speech distinctions are recoverable from SAE activations via compact latent groups, not one-to-one latent-to-category mappings.
Read by the desk
↑10hf upvotes
RESEARCHQwen-Planner-Agent: A Closed-Loop AI-for-AI Framework for Real-World Mobile Planner Agents · arXiv · Sep 26
RewardVerse uses generated rubrics as an intermediate step before scoring to reduce scalar drift in video reward models trained with a two-stage RL algorithm.
Read by the desk
↑22hf upvotes
RESEARCHSchrödinger's Code Repository: Have LLMs Learned SWE-bench or Memorized It? · arXiv · Sep 25
An evaluation framework that dynamically transforms test repositories to detect whether coding agents rely on memorized cues rather than genuine reasoning.
MemBodied adds fixed-size episodic memory to Vision-Language-Action models, achieving 7.81× the success rate of a stateless policy on memory-dependent tasks.
GAE is an autoencoder whose latent space decodes jointly to appearance, depth, cameras, and point maps, cutting FVD by 12.7% and 23.1% on two benchmarks.
Read by the desk
↑43hf upvotes
RESEARCHAll-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts · arXiv · Sep 24