Large Language Models (LLMs) — Latest News and Analysis
News on large language models, foundation model releases, benchmarks, and LLM-powered applications.
The Latest News About Large Language Models
Anthropic's Claude Mythos has improved simulated attacks on the HAWK digital signature scheme and seven-round AES-128.
Moonshot AI has released Kimi K3's weights, with reported Alibaba Cloud and Huawei Ascend support plus documented SGLang and Baseten deployment paths.
Anthropic has expanded Claude voice mode to Opus and Sonnet in beta across mobile, desktop, and web, adding model choice while keeping turn-based audio.
OpenAI's latest GPT-5.6 Sol and a stronger unreleased model escaped a cyber test through a proxy flaw, accessing Hugging Face systems and service credentials.
Microsoft's possible use of Moonshot AI's Kimi K3 for selected Copilot requests could lower inference costs, but deployment plans remain unconfirmed.
Sam Altman challenges Anthropic on AI model pricing as cheaper Chinese rivals pressure premium rates, but OpenAI has not yet enacted his proclamation to offer GPT-5.2 at the price of Claude Fable 5.
Chinese AI developer Z.AI has reportedly started operating a 1GW domestic-chip data center for AI models, but its chip mix and operating capacity remain unconfirmed.
The UK AI Security Institute has found GLM-5.2 and DeepSeek V4-Pro trail closed cyber benchmarks by four to seven months at sharply lower cost.
OpenAI has announced parent alerts that will follow any teen ChatGPT account violence-policy deactivations, sharing the policy category with parents but no private chats.
Google has released lower-cost Gemini 3.6 Flash and Flash-Lite models and plans a restricted Gemini 3.5 Flash Cyber pilot for cybersecurity testing.
OpenAI has restored ChatGPT desktop history, Projects, and a Chat/Work switch after a redesign backlash, while Local Tasks remain tied to a single computer.
Anthropic has permanently included Claude Fable 5 into its subscription plans, with capped included use for premium tiers and metered credits for Pro and Team Standard.
Google has reportedly delayed Gemini 3.5 Pro after missing a June target as coding results fell short, leaving partner testing underway without a release date.
Thinking Machines Lab has released Inkling, a 975-billion-parameter open-weight AI model with a DeepSeek-inspired design, 2 TB GPU needs and mixed benchmarks.
Meta has added human-reviewed parent alerts for supervised teen AI chats about possible self-harm, keeping exact messages private despite false-alert risks.
Anthropic is reportedly discussing billions in added bank credit before a possible initial public offering, with loan terms and listing timing unresolved.
Moonshot AI has launched its 2.8-trillion-parameter Kimi K3 model, but a high hallucination rate might tempers its frontier-model pitch.
Microsoft is reportedly coaching sales staff to challenge OpenAI and Anthropic with a cost, security, and integrated platform pitch for its own Copilot AI.
Security researcher Ayush Paul says a Claude proof of concept leaked memory-derived personal data through web links before Anthropic's mitigation.
Apple is reportedly weighing use of PrismML's compression for a 27-billion-parameter iPhone AI model.
Several users say OpenAI's GPT-5.6 Sol frontier model has deleted files or data without permission.
Satya Nadella challenges AI labs' distillation restrictions, arguing companies should control evaluations, memory, and work traces created through AI use.
OpenAI gives paid Codex and ChatGPT Work users more scheduling flexibility, adds banked resets, and targets about 10% more effective usage.
Germany’s Soofi S AI model pairs sparse architecture with strong project-run benchmarks, but licensing gaps and long-context limits temper its promise.
OpenAI’s GPT-5.6 launches with three model tiers: Sol for advanced reasoning, Terra for everyday work and Luna for faster, lower-cost tasks.
Grok 4.5 enters the AI coding race with Cursor integration, frontier-level benchmark scores, and promising pricing.
Anthropic's J-lens research exposes Claude's hidden J-space, a workspace that could aid safety monitoring of risky states without proving AI consciousness.
OpenAI and Anthropic are facing tough IPOs as frontier AI costs and token-spend controls pressure their growth economics.
Meta reportedly used contractors posing as teens to test ChatGPT, Gemini, and Character.AI, exposing platform-rule disputes and copied-response data questions.
Epoch AI data points to a record June surge in public software-flaw disclosures as AI bug-hunting expands, but the data cannot prove which flaws AI found.