Daily News — 2026-08-19
Digest #1 · new work yesterday
-
HarnessEval-W: Agentifying the Evaluation of Visual Worlds
A benchmark framework that uses hierarchical sub-agents to decompose world-model evaluations into verifiable reasoning chains with transparent evidence.
-
Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization
A policy optimization method that standardizes multi-objective rewards and discounts saturated objectives to focus training on under-optimized goals.
-
Large Discovery Models: Empirically-grounded Model-Based Open-Ended Search
A recurrent Large Discovery Model pairs generative proposals with a Bayesian non-parametric reward surrogate, tested across molecules, proteins, and programs.
-
VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End?
RL training lets open-source multimodal agents plan and build 3D scenes end-to-end, outperforming closed-source frontier models on this benchmark.
-
MOSS-VL Technical Report
MOSS-VL is an open vision-language model family that cuts time-to-first-token latency for real-time streaming interaction via gated cross-attention.
-
UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations
A foundation GUI agent that combines environment-grounded training with in-context demonstration learning to improve reliability on long-horizon office tasks.
-
ClawGym II: Exploring Black-Box RL on Agent Harness
A black-box RL framework for optimizing agents through complex harnesses via sandbox execution, trajectory reconstruction, and mix-harness training.
-
An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models
Latent-to-pixel transfer training cuts convergence time and inference cost for large-scale pixel-space diffusion models.
-
Agentic Transaction: Towards ACID-Compliant Agent Systems
A framework that brings ACID-style semantic guarantees to long-horizon LLM agent workflows.
-
ifixai-ai/iFixAi
A tool for auditing AI agents to verify they are doing what they are supposed to do, returning results in under 120 seconds.
-
internet-court/internet-court-skill
Trust layer for agent-to-agent commerce: natural-language mandates, ERC-7710 permissions, x402 payments, escrow, and dispute resolution in one open skill.
-
stablyai/orca
Orca is a desktop and mobile ADE for running a fleet of parallel coding agents with your own subscription.
-
tt-a1i/archify
Agent skill producing self-contained HTML diagrams—architecture, workflow, sequence, data-flow, and lifecycle—with motion and crisp export.
-
holaboss-ai/holaOS
An open-source AI agent workspace that runs agents across tools, apps, browser, and files with shared memory and 100+ integrations.
-
akitaonrails/ai-memory
A repository providing long-term memory for agent coding CLIs and handoff between different agent vendors.
-
coreyhaines31/marketingskills
Plug-in marketing skills for Claude Code agents: CRO, copywriting, SEO, analytics, and growth engineering in one repo.