AI Agents are rapidly evolving from tools into a unified interface for user interaction—ushering in the 'Super Assistant' paradigm that replaces traditional app ecosystems. Meanwhile, Zhipu AI has become the world's highest-valued open-source software company by market capitalization, surpassing Xiaomi [4]. Thought leaders like Naval Ravikant emphasize embracing 'irrational optimism' to navigate AI-driven systemic transformation amid organizational restructuring and a hardware renaissance [3].
Start with the newest briefing, then Continue by task
The newest briefing gives you today's main changes in a few minutes. Use the focused routes for evidence, implementation detail, and longer analysis.
Posts
Embodied world models are experiencing an explosive wave of open-source progress, with τ0-WM and STI-WM recently released—marking a new phase for robots' 'slow thinking' decision-making and the real-world deployment of physical AI. Zhipu AI has surged to become the world's highest-valued open-source software company through its full-model open-source strategy, with a valuation exceeding Xiaomi's [1]. Meanwhile, Anthropic is reportedly suspected of deliberately degrading older model performance—a revelation sparking deep ethical concerns regarding commercial practices among large-model vendors [10].
Anthropic's valuation has surged to $96.5 billion, officially surpassing OpenAI to become the world's highest-valued AI company; meanwhile, general-purpose AI Agents are being defined by multiple experts as the 'next-generation operating system,' rapidly reshaping app paradigms, SaaS architectures, and enterprise organizational models [8][2][4].
Anthropic missed GUI product opportunities due to overreliance on the TUI interaction paradigm—highlighting the design advantages of the Claude App; meanwhile, China's first full-stack green-AI computing platform launched in Inner Mongolia, integrating compute orchestration, model invocation, and token trading—marking a new phase of low-carbon, collaborative AI infrastructure [1][7].
Microsoft releases 45-year-old MS-DOS source code, restored via OCR from paper archives; Xiaomi unveils MiMo-V2.5—its full-stack inference optimization architecture with Hybrid SWA, MoE, and multimodal synergy; Fireworks AI's valuation hits $1.5B, signaling rapid capitalization of AI inference infrastructure.
AI agents are shifting from Copilot-style assistance to autonomous SDLC execution—Salesforce cut key workflows from 231 to 13 person-days. Meanwhile, Gamma-World, a multi-agent world model, overcomes identity symmetry and communication bottlenecks—advancing embodied AI architecture.
Claude Opus 4.8 introduces mid-conversation system messages, significantly enhancing agent controllability and engineering robustness [10]; BYD launches the Juxuan A3—a vehicle-grade, in-house 4nm AI chip—matching NVIDIA in compute and energy efficiency, signaling a new phase in China's physical AI hardware competition [12]; Global memory export prices surge nearly 1,000%, reflecting supply-chain restructuring pressures amid explosive demand for AI compute infrastructure [11].
China's AI industry is exhibiting a paradoxical 'high-investment, low-perception' divide: On one side, Qijing GT7 integrates Huawei's Qwen Intelligent Driving and HarmonyOS Cockpit for full-stack technical integration; on the other, global tech giants like Amazon have explicitly drawn red lines against AI misuse. Meanwhile, Grok Build 0.1 has officially launched in Cursor—the developer's primary IDE—and industry-wide reflection is intensifying [1][2][3][4].
Claude Opus 4.8 has officially launched, significantly enhancing programming capabilities and dynamic Subagent workflow support—enabling concurrent orchestration of hundreds of sub-agents for complex tasks. Meanwhile, a Tsinghua-affiliated team's breakthrough 'Intelligent Compute Grid' technology is overcoming bottlenecks in deploying domestic AI compute: via heterogeneous pooling, it transforms domestically developed chips into highly available, low-cost, standardized token production capacity [13][11].
The new Claude Code `/usage` command launches—marking the first production-grade, token-level granular tracking of consumption across four agent capability types: Skills, Agents, MCPs, and Plugins—ushering AI engineering into the era of 'measurable cost'.
Anthropic is redefining the boundaries of agent capabilities with Claude Opus 4.8 and Dynamic Workflows—while industry consensus rapidly shifts toward recognizing that an agent's true capability lies in its accessible tools and execution scope, not anthropomorphic role-playing. Meanwhile, major tech companies are broadly trapped in a 'black-box accounting' dilemma: unclear ROI and runaway AI budgets [1][2][5].
China's AI infrastructure ecosystem is shifting toward chip-model co-design, highlighted by DeepSeek V4 and the Kunpeng-Ascend Summit; meanwhile, Claude Code's cloud deployment is gaining traction, with Alibaba ATA and community guides advancing it toward production-ready multi-user, streaming, and sandboxed architectures.
The domestic large language model Qwen3.7 Max ranked second globally in real-world Vibe Coding (atmospheric programming) benchmarking—outperforming several leading international models; meanwhile, SK Hynix's market capitalization surpassed USD 1 trillion, making it the world's first memory chip manufacturer to join the 'Trillion-Dollar Semiconductor Club' [1][2].
Agent engineering is rapidly transitioning from proof-of-concept to production deployment: Alook's open-source platform enables role-based orchestration of CLI Agents; Fudan NLP's team offers a fully automated research Agent for academia—complete with free GPU support. Meanwhile, Xiaomi has slashed its domestic large-model API pricing to ¥0.025 per million tokens, signaling the industry's deep dive into an intense 'token price war' [16].
The ultimate form of mobile AI is evolving toward 'imperceptible intelligence'—OPPO's ColorOS 16 and vivo's official website AI shopping assistant both validate a lightweight deployment path combining compact intent-recognition models, Agent workflows, and RAG knowledge bases [0][5]. Meanwhile, AI commercialization faces structural bottlenecks: both advertising and subscription models have hit saturation points, and industry consensus is rapidly shifting toward an 'execution economy' centered on 'task completion' [9].
Tactile embodied AI secures ~$10M angel round; OpenRouter raises $113M in Series B, hitting 25T weekly tokens; SynthID has watermarked 100B AI-generated items and is integrating into Google Search & Chrome.
CUDA 13.3 officially introduces C++ Tile programming and the CompileIQ auto-tuning framework—marking a paradigm shift toward higher-abstraction GPU development. Meanwhile, Stack Overflow—despite a steep decline in user post volume—achieved $115 million in annual revenue through its enterprise AI knowledge base and data licensing services, validating a new commercialization path for developer platforms in the AI era [2][3].
AI engineering is rapidly entering a new era of 'AI building AI': Baichuan Intelligence has launched ForgeTrain—the world's first production-grade pretraining framework fully written by AI—and successfully trained MiniCPM5-1B. Concurrently, foundational architectural innovations—including Domain-Specific Architectures (DSAs), KV cache quantization (e.g., OSCAR's 2-bit scheme), and the newly proposed 'Tao Law'—are being deployed at scale, continuously pushing past compute and energy-efficiency bottlenecks [24][10][22].
Huawei proposes a new chip evolution paradigm—'The Tao Law'—centered on the time constant τ, challenging the traditional Moore's Law trajectory; meanwhile, DeepSeek tops the global large-model API call leaderboard, highlighting the scalable deployment of domestic AI infrastructure [3]; and 'commitment hallucination' in AI conversational products—triggered by anthropomorphic design—is exposing deep gaps between product accountability and legal regulation [0].
AI engineering is rapidly evolving from prompt engineering toward framework engineering and context engineering; standardization of Agent Harness and vertical-domain workflow reengineering have become critical for real-world adoption. DeepSeek has entered the programming-agent market with a 'Mixue Bingcheng'-style low-cost strategy—directly positioning itself against Claude Code [2][6][7][2].
DeepSeek launches a low-price offensive with its permanently discounted V4-Pro API, targeting the Claude Code–level programming agent market; Baishan Intelligence breaks the edge-side bottleneck by achieving 1.58-bit ternary quantization training of a 60-billion-parameter model on Huawei's Ascend platform—reducing GPU memory usage by 6× while retaining 97% of model capability [2]; meanwhile, Gemini quietly overhauls its billing logic, drastically shrinking actual usage quotas for paying users and exposing a trust gap in large-model commercialization [10].
Baobei Intelligence, Tsinghua University, and OpenBMB achieved end-to-end training of a 60B-parameter LLM on Huawei Ascend using 1.58-bit ternary quantization—cutting memory use by ~6× while retaining 97% capability. Meanwhile, continuous-space language modeling emerges as a paradigm shift beyond token-based autoregression, seen as a key step toward AGI.
Agent tech is maturing rapidly—Codex and similar tools are enhancing core workflow capabilities. Meanwhile, Google's CEO acknowledged Gemini's gaps in coding agents and long-horizon tasks, signaling a shift from model benchmarks to real-world task completion. Anthropic's 'should do' > 'can do' framework highlights the growing scarcity of AI judgment.
AI is accelerating into the agent-native era—coding capability is now the key differentiator for agents [5]; AI compute is shifting historically toward inference, expected to consume 70% of total AI compute [17]; Anthropic reveals Claude's new 'dreaming' memory and long-horizon collaboration architecture [9], while Google's CEO admits Gemini lags in coding agents [16].
AI industry focus is shifting structurally: inference compute will rise to 70% of total AI spend; Agent startups now compete for enterprise payroll budgets—not just SaaS budgets—while high-quality medical data and structured security skill repositories emerge as key new barriers.