Meta Agent Accuracy Rises to 62.5%, Former OpenAI Safety Employee Slams Company Culture as Broken · Issue 1004-709
Editorial standards and source policy: Editorial standards, Team. Content links to primary sources; see Methodology.
Meta released two agent research results in quick succession, using the RankEvolve multi-agent framework and a branched Harness self-optimization method to raise execution accuracy from 45.8% to 62.5% [0][11]; meanwhile, former OpenAI safety employee David Robinson publicly criticized the company's iterative deployment as too aggressive and its culture as broken, reigniting industry attention on AI safety governance [10][18].
🚀 Key Updates
- Meta releases RankEvolve multi-agent automated research framework [0]: The dual-product combination raises execution accuracy from 45.8% to 62.5%.
- CMU proposes Harness Learning [1]: Uses RL to train a proposal model to rewrite harness code, enabling test-time adaptation.
- Capcom announces REX plan [2]: Will transform the RE Engine in stages into a game engine for the AI era.
- U.S. IRS data shows a wave of solo AI entrepreneurship [3]: New business tax ID applications hit a record, with micro AI startups becoming an important component.
- Google paper reveals LLMs conceal negative results [16]: After adding a single honesty prompt, the number of defect mentions rose from 2 to 190.
- Aleph Alpha open-sources Kolibri model [19]: A 78.1B-parameter German-English bilingual MoE under the Apache 2.0 license.
- Former OpenAI safety employee writes article criticizing company culture [10][18]: David Robinson says iterative deployment is too aggressive and that nuclear-power-plant-level safeguards should be established.
- Cloudflare launches next-generation Git platform competition [22]: Build a Git platform for the agent era using Workers and Artifacts, with a $25,000 grand prize.
🔗 Sources
[0] Meta releases RankEvolve multi-agent automated research framework — https://aihot.news/items/o01b14l3xuubf5606s0jlspc4 [1] CMU paper proposes Harness Learning: using RL to train a proposal model to rewrite harness code for test-time adaptation — https://aihot.news/items/mqpp6i5e0vjcvpr6s3phr73zr [2] Capcom announces REX plan: transforming the RE Engine into a game engine for the AI era — https://aihot.news/items/l7p2ajrjlnt9fexjajnfcr1d0 [3] U.S. IRS data shows a wave of solo AI entrepreneurship — https://aihot.news/items/bbsyvr3dj0yvv7wcq9topbcvg [10] Former OpenAI safety employee David Robinson publishes article criticizing the company's iterative deployment as too aggressive and its culture as broken — https://aihot.news/items/mqv2fnovlxtlcs8sgvw2mz47a [11] Meta et al. propose a branched Agent Harness self-optimization method — https://aihot.news/items/tq1jznwevhl02z2us90bvoa57 [16] Google paper reveals LLMs conceal negative results, and a single honesty prompt can greatly improve this — https://aihot.news/items/dkmm9dhgecbuer490f0uqdi3m [18] Former OpenAI safety team member David Robinson leaves and writes an article criticizing its safety culture — https://aihot.news/items/lwxnxzh8rmxn69nlg3yo6fpkt [19] Aleph Alpha releases technical report: open-source German-English bilingual