Updates

Official digests and analysis

Start with the newest briefing, then Continue by task

The newest briefing gives you today's main changes in a few minutes. Use the focused routes for evidence, implementation detail, and longer analysis.

Posts

Claude Code Turns One: 40x Memory Reduction + Cross-Device Remote Support · 0225-60

Claude Code achieves dual breakthroughs on its first anniversary: p99 memory usage drops by 40×, and cross-device Remote Control officially launches; meanwhile, industry consensus rapidly converges—shifting decisively from 'programming for humans' to 'building for AI Agents', with CLI, observability, and outer-loop closure as foundational infrastructure.

Rolling Sink: Bridging the Temporal Gap in Video Diffusion · 0225-58

The temporal gap in video diffusion models is being systematically bridged by the Rolling Sink mechanism; Anthropic accelerates enterprise AI collaboration with Claude Cowork and an industry-specific plugin matrix; Qdrant 1.17 introduces native relevance feedback for vector indexes—the first of its kind—redefining production-grade search optimization; Meta and AMD have signed a multi-year agreement to deeply integrate AMD Instinct GPUs into Meta's planned 6GW AI data center infrastructure, underscoring a strategic upgrade in compute infrastructure.

Anthropic Accuses DeepSeek and Two Other Companies of Industrial-Scale Model Distillation · 0224-57

Anthropic publicly accused DeepSeek, Moonshot AI, and MiniMax of conducting 'industrial-scale distillation attacks,' sparking broad debate on AI model security and intellectual property boundaries; meanwhile, the industry is accelerating its shift toward AI Agent engineering—exemplified by OpenAI's Codex App, TinyFish's $2M seed fund for AI Agents, and the Claude Code + Obsidian personal operating system.

OpenAI Releases GPT-Realtime-1.5 Model, Cutting Time-to-First-Token by 40% · 0224-56

OpenAI overhauls real-time capabilities: launching the gpt-realtime-1.5 model and adding WebSocket support to the Responses API—cutting time-to-first-token (TTFT) by up to 40%. Meanwhile, Anthropic introduces a novel 'Persona Selection Model' to explain Claude's human-like behavior—and accuses several Chinese labs of large-scale 'distillation attacks'.

Anthropic Releases AI Fluency Index · 0224-55

Anthropic officially launches the 'AI Fluency Index,' redefining human-AI collaboration assessment through 11 collaborative behaviors; meanwhile, Llama 3.1 8B achieves inference speeds exceeding 18,000 tokens/sec—pushing the performance frontier of on-device AI via hardware-level parameter hardening.

Llama 3.1 8B Inference at 18,000 Tokens/Second · 0223-54

AI inference performance achieves a hardware-level breakthrough—Llama 3.1 8B reaches 18,000 tokens/sec; meanwhile, GLM-5 achieves full-stack compatibility with domestic chips, and the COMI framework outperforms baselines by 25 points under 32× long-context compression—signaling dual leaps in model efficiency and indigenous capability...

GLM-5 and Antigravity Drive AI System-Level Restructuring, Accelerating the Erosion of SaaS Moats · 0223-53

AI is rapidly evolving—from agent engineering (e.g., GLM-5, Antigravity) toward system-level rearchitecture: 'File System as Database,' 'Code as Tool' (MCP architecture), and 'Sketch as Application' are emerging as new paradigms; meanwhile, SaaS moats continue to erode, confirming that AI is fundamentally redefining software complexity and commercial barriers.

GLM-5 Officially Released: DSA Sparse Attention + Asynchronous Reinforcement Learning for Agent Engineering · 0223-52

At the start of 2026, U.S.-China AI development has entered a high-frequency race—30 major updates in just 47 days; GLM-5 has officially launched, advancing AI toward the new paradigm of 'Agent Engineering' via DSA sparse attention and an asynchronous reinforcement learning infrastructure; Beijing's Haidian District has emerged as the strongest hub for breakthroughs across all modalities and the full AI industry chain.

Gemini 3.1 Pro Converts Academic Papers into Executable Code · 0222-50

Gemini 3.1 Pro demonstrates remarkable capability in directly converting cutting-edge academic papers (e.g., Local-First CRDT) into runnable simulation programs; meanwhile, OpenAI’s Batch API now supports GPT image models for the first time—reducing batch task costs by 50%, marking a milestone in multimodal scaling...

Taalas HC1 Chip Delivers 17,000 Tokens/s Inference Performance · 0222-49

AI infrastructure is undergoing a dual shock: an ASIC hardware revolution and a precipitous drop in inference costs. The Taalas HC1 chip delivers 17,000 tokens/sec inference throughput at just $0.0075 per million tokens; meanwhile, NVIDIA has shifted to strategic capital alignment—investing $30 billion directly into OpenAI, marking its evolution from a 'pick-and-shovel' supplier to a co-builder.

Taalas Unveils 17,000-Token/s ASIC to Challenge NVIDIA · 0221-48

AI hardware and software stacks are undergoing simultaneous, accelerated redefinition: Taalas challenges NVIDIA's compute dominance with a purpose-built ASIC chip delivering 17,000 tokens per second, while NVIDIA pivots to strategic capital alignment—investing $3 billion directly into OpenAI. Meanwhile, Claude Code undergoes a comprehensive upgrade in agent collaboration capabilities, and its new Git Worktree support plus non-Git system compatibility signal that AI-powered programming infrastructure has entered a deep, production-grade engineering phase.

Gemini 3.1 Pro, Lyria 3, and Claude Code Launch This Week · 0221-47

Gemini 3.1 Pro, Lyria 3, and Claude Code form this week's 'trident' of AI engineering advancement: Google strengthens systematic engineering reasoning and multimodal creation capabilities, while Anthropic accelerates the shift toward AI-native development with a 1M-token context window, a research preview of Code Security, and cross-platform conversation migration.

Integrate Llama.cpp with Hugging Face · 0221-46

Llama.cpp has officially integrated into the Hugging Face ecosystem—signaling deep synergy between lightweight inference and model distribution infrastructure; GPT-5.2 Thinking outperforms Gemini 3 DeepThink on world-knowledge reasoning tasks, underscoring 'chain-of-thought depth' as a critical differentiator for next-generation large language models.

Gemini 3.1 Pro's Logical Reasoning Soars to 77.1%, Far Outpacing Its Predecessor · 0220-45

Gemini 3.1 Pro has officially launched, with its logical reasoning performance on the ARC-AGI-2 benchmark surging to 77.1% (up from just 31% in the previous version), outperforming competitors across multiple metrics. Meanwhile, OpenAI CEO Greg Brockman explicitly identified reasoning compute as the current core bottleneck for software productivity...

Gemini 3.1 Pro Tops Multidimensional Benchmarks, Doubles Logical Reasoning Performance · 0220-44

Gemini 3.1 Pro has officially taken the top spot across multidimensional benchmark suites, doubling its logical reasoning capability (achieving 77.1% on ARC-AGI-2) and propelling Google back into the AI model vanguard; meanwhile, OpenAI, Perplexity, Replit, and Anthropic—among other industry leaders—are rapidly upgrading interaction paradigms—from real-time Mermaid previews and direct SEC filing audits to automatic prompt caching—ushering AI development and usage into a new era of 'what-you-see-is-what-you-get' and 'verifiably trustworthy' experiences.

NTT DATA Boosts Japanese AI Accuracy from 15.3% Using Nemotron-Japan Dataset · 0220-43

A pivotal breakthrough in Japanese-language AI deployment: NTT DATA leveraged NVIDIA's Nemotron-Personas-Japan synthetic dataset to boost model accuracy from 15.3% to 79.3%; meanwhile, Anthropic tightened ecosystem permissions—fully disabling OAuth integration—highlighting large-model vendors' dual emphasis on security and control.

Claude Opus 4.6 Launches with 1-Million-Token Context Window · 0219-42

AI is rapidly evolving beyond the tool layer into the decision-making layer: Claude Opus 4.6 redefines capability boundaries with its 1-million-token context window and dynamic computation; domestic large models—including Ling-2.5-1T and Qwen3.5-397B-A17B—have surged into the global top tier of open-source LLMs; meanwhile, distribution capabilities and Agent security architecture have replaced coding efficiency as the new bottleneck—and decisive battleground—for growth.

Qwen3.5 Full-Stack Optimization Unleashed: NVIDIA & AMD Support · 0218-41

The Qwen 3.5 series—including the 397B-A17B and Plus variants—is triggering explosive, full-stack ecosystem adoption across leading hardware platforms and developer toolchains—from NVIDIA NeMo and AMD Instinct GPUs to Ollama Cloud, ZenMux, and mlx-vlm—with first-day support now live. Meanwhile, LlamaIndex is accelerating its evolution toward a token economy, restructuring API access around the $LLAMA token.

Qwen 3.5 Launches and Is Immediately Optimized for NVIDIA GPUs · 0218-40

The Qwen 3.5 series is triggering a full-stack ecosystem surge—major hardware vendors including NVIDIA and AMD, as well as development platforms such as Ollama Cloud, ZenMux, and mlx-vlm, have all delivered day-one support. Meanwhile, LlamaIndex is rapidly evolving into foundational AI Agent infrastructure—redefining its API economy via the $LLAMA token and enhancing multimodal data processing with LlamaCloud's advanced PDF parsing.

Qwen3.5 Released: 397B-Parameter Native Multimodal Model, First-Day Full-Stack Support from NVIDIA and Others · 0217-39

The Qwen 3.5 series has powerfully ignited the open-source LLM ecosystem—its 397B-parameter count, native multimodality, and MoE + Linear Attention architecture received full-stack Day-One support from NVIDIA, AMD, Ollama, ZenMux, LMSYS, and mlx-vlm; meanwhile, LlamaIndex accelerates its evolution into AI Agent infrastructure—replacing subscriptions with the $LLAMA token and upgrading PDF-to-Markdown/JSON parsing capabilities to strengthen agents' 'cognitive infrastructure.'

Qwen3.5-397B Open-Source LLM Enhances B2B Practical Capabilities · 0217-38

AI is rapidly shifting from capability augmentation to role replacement: LLM-powered code translation, visual UI editing, and memory-driven agents are emerging as new productivity foundations; open-source large models like Qwen3.5-397B continue strengthening B2B operational capabilities, while teams led by Fu Sheng and Google's Antigravity project independently validate the scalable real-world deployment of AI assistants in personalized content distribution and human-AI collaborative editing.

Alibaba Open-Sources Qwen3.5-397B-A17B: The World's First Native Multimodal Sparse MoE LLM · 0217-37

Alibaba officially open-sourced Qwen3.5-397B-A17B—the world's first natively multimodal, sparse Mixture-of-Experts (MoE) large language model, supporting 1M-token ultra-long context and 4-bit local inference on consumer-grade hardware. Meanwhile, Manus Agents launched long-term memory and toolchain integration on Telegram—marking AI assistants' evolution into the 'memorable and actionable' era.

OpenAI Poaches OpenClaw Founder Amid MiniMax's Valuation Surge · 0216-36

OpenAI is strategically accelerating its expansion into the personal agent ecosystem, notably recruiting Peter Steinberger, founder of OpenClaw. Meanwhile, MiniMax has achieved a valuation leap through its highly cost-efficient, reasoning-optimized technical approach, and Claude Code is now supporting an annualized $2.5 billion...