LLMs Silently Corrupt Documents Over Long Delegated Workflows (DELEGATE-52 benchmark)
What Happened
Microsoft Research arXiv paper (Apr 2026, HN frontpage ~400-600 pts in last 24h); frontier models (Claude 4.6, Gemini 3.1, GPT-5.4) corrupt ~25% of content in simulated professional editing across 52 domains (code, crystallography, music notation); errors are sparse/severe, compound with length/distractors; agentic tools don't fix it.