AI Agents — Latest News and Frameworks
Coverage of autonomous AI agents, frameworks, and the shift toward agentic AI workflows.
The Latest News About AI Agents
Anthropic's Claude Opus 5 offers near-Fable prformance at Opus-tier pricing, but benchmark claims, safety routing and enterprise ROI still need validation.
OpenAI's latest GPT-5.6 Sol and a stronger unreleased model escaped a cyber test through a proxy flaw, accessing Hugging Face systems and service credentials.
Jack Dorsey's Block has launched Buzz to combine team chat, AI agents and Git work, though the early-stage workspace relies on one authoritative relay.
The UK AI Security Institute has found GLM-5.2 and DeepSeek V4-Pro trail closed cyber benchmarks by four to seven months at sharply lower cost.
OpenAI has restored ChatGPT desktop history, Projects, and a Chat/Work switch after a redesign backlash, while Local Tasks remain tied to a single computer.
Microsoft reportedly plans Project Perception, a multi-model AI bug finder, but its July timing, design, availability and performance remain unconfirmed.
Microsoft made Copilot Cowork available worldwide on June 16, pairing completed multi-tool work with usage billing and direct administrator cost controls.
Sakana AI has outlined plans to add Nvidia Nemotron specialists to its Fugu AI orchestration platform.
Microsoft is consolidating its security-engineering roles with several hundred layoffs as teams shift toward AI defenses and Security Copilot work.
SpaceXAI has opened Grok Build's coding-agent harness for inspection and local use after hidden uploads of complete codebases by the tool were discovered.
Blume turns Markdown folders into static Astro documentation and AI-readable files.
DeepSeek is seeking a $7.4 billion raise at a contemplated $74 billion valuation after a recent June round, but terms and futureIPO listing plans remain preliminary.
OpenAI's Codex agent is now encrypting messages between agents, resulting in unreadable local task records that can hinder audits and debugging after handoffs.
Several users say OpenAI's GPT-5.6 Sol frontier model has deleted files or data without permission.
OpenAI gives paid Codex and ChatGPT Work users more scheduling flexibility, adds banked resets, and targets about 10% more effective usage.
OpenAI’s GPT-5.6 launches with three model tiers: Sol for advanced reasoning, Terra for everyday work and Luna for faster, lower-cost tasks.
Cognition has launched SWE-1.7 for its Devin software engineering agent with near-frontier coding scores.
Grok 4.5 enters the AI coding race with Cursor integration, frontier-level benchmark scores, and promising pricing.
A look at GitLost, the GitHub Agentic Workflows prompt-injection case where public issue text, private repo access, and public comments collide.
Former GitHub-CEO's startup Entire has launched a Git network that mirrors GitHub repositories for AI coding agents with regional cells.
Anthropic has expanded Claude Cowork to mobile and web, widening access beyond the desktop app while desktop remains the fuller work surface for local tasks.
OfficeCLI offers command-line tools for AI agents to work efficiently with Word, Excel and PowerPoint files.
Sysdig says JADEPUFFER may be the first agentic ransomware case, exposing how AI agents can turn old credential failures into database destruction.
Alibaba is banning employee the use of Anthropic's Claude Code starting July 10 and shifts staff to its own Qoder platform over alleged user-identification security risks.
Alibaba Cloud's SkillWeaver framework routes AI-agent tasks to relevant tools and claims 99% lower benchmark token use, but code and production proof remain open.
Microsoft is reportedly working on a unified Copilot app that could merge consumer and enterprise apps, cut underused features, and make the AI assistant prove practical value to users.
Google rolls out Gemini Spark on Mac for eligible Google AI Ultra subscribers, adding permission-based file automation, app links, and staged remote tasks.
Anthropic is reportedly preparing Claude for Microsoft Teams, testing how workplace agents handle channel access, tools, billing and governance controls.
Anthropic has launched Claude Sonnet 5 for lower-cost multi-step AI agent work, with broad developer access, dicounted release API pricing, and tokenizer caveats.
Google Cloud is positioning its Spanner SQL database service as a multi-model database for AI agents, combining graph, vector, analytics and hybrid deployments.