Latest AI News

Last 72 hours, organized by category

🟡 🤝 Agents July 10, 2026 · 2 min read

LangChain: OpenWiki Brains gives agents proactive memory — automatically collects context from Gmail, Notion and Git into Markdown

Editorial illustration: a robot assembling a wiki from Markdown leaves pulled from email, notes and repositories

OpenWiki Brains is LangChain's open-source framework that gives AI agents proactive memory: it automatically collects context from external sources and saves it as Markdown files on disk. The first version supports 6 connectors — Gmail, Notion, Git, Twitter/X, Hacker News and web search — and runs locally through scheduled jobs instead of relying on the cloud.

🟢 🤝 Agents July 10, 2026 · 2 min read

arXiv:2607.08403: Game Theory Against Hallucinations — Multi-Agent Framework Models LLMs as Strategic Players to Reduce Confabulation

Editorial illustration: several AI players at a chessboard coordinating their moves toward the truth

A new paper proposes a multi-agent framework using game theory to reduce hallucinations in large language models. Model interactions are modeled as a game of strategic players where equilibrium corresponds to consistent, non-hallucinated responses — offering a path distinct from standard RLHF and RAG approaches.

🟡 🤝 Agents July 9, 2026 · 2 min read

arXiv:2607.06906: 'Harness Effect' — orchestration design cuts AI agent cost by 41%, more than changing the model

Editorial illustration: a conductor's baton directing token flows between models and tools

The Harness Effect is a finding from an empirical paper by 32 authors showing that the design of the orchestration layer affects AI agent costs more than the choice of the model itself. Across 22 tasks and 6 models, optimized orchestration reduced cost per task by 41% (from $0.21 to $0.12), tokens by 38%, and execution time by 44% — while also improving quality.

🟡 🤝 Agents July 9, 2026 · 2 min read

IBM: Bob becomes a multi-agent platform — nine-month COBOL modernization project completed in 3 days

Editorial illustration: mainframe cabinet from which a swarm of small agents transfers code to modern servers

IBM Bob is an agentic platform for software development that in its new version gains a multi-agent architecture for the full development lifecycle, Bobalytics analytics, and a Premium package for mainframe modernization of COBOL and PL/I. IBM reports that client Blue Pearl completed a nine-month modernization project with Bob in just three days.

🟡 🤝 Agents July 9, 2026 · 2 min read

OpenAI: ChatGPT becomes an agent for serious work — operates for hours through apps and files until the job is done

Editorial illustration: a work desk with multiple screens through which an automated task flow runs

ChatGPT Work is OpenAI's move into agentic work: ChatGPT can now take actions across apps and files and work independently on a project for hours, turning a stated goal into completed work. OpenAI positions it as a partner for the most ambitious tasks, competing directly with agentic products from Anthropic and GitHub.

🟡 🤖 Models July 10, 2026 · 2 min read

arXiv:2607.08733: 'Super Weights' Explain Why Selective Fine-Tuning Fails — Paper Accepted at COLM 2026

Editorial illustration: a network of parameters in which a few bright nodes hold the entire structure together

Super Weights in LLMs is a paper accepted at COLM 2026 that identifies 'super weights' — a small number of parameters whose modification disproportionately changes a language model's behavior. The paper shows that these weights explain why selective fine-tuning of only certain layers often fails, with direct implications for PEFT and LoRA approaches.

🔴 🤖 Models July 9, 2026 · 2 min read

Google: SensorFM — foundation model trained on one trillion minutes of wearable data wins on 34 of 35 health tasks

Editorial illustration: smartwatch with waves of biometric signals flowing into a neural network

SensorFM is Google's foundation model for health data from wearable devices, trained on more than one trillion minutes of signals from Fitbit and Pixel Watch devices worn by 5 million users in over 100 countries. The model outperforms specialized approaches on 34 of 35 tasks, with +9% AUC on classification and +21% correlation on regression.

🔴 🤖 Models July 9, 2026 · 2 min read

Microsoft: Aurora 1.5 open-source model outperforms ECMWF ensemble on 88.9% of variables — new standard in AI weather forecasting

Editorial illustration: globe with swirling weather fronts and data network

Aurora 1.5 is Microsoft's open-source foundation model for the Earth system that outperforms the ECMWF ensemble forecast on 88.9% of evaluated variables and horizons. The new release adds 22 meteorological variables, hourly resolution, and probabilistic ensemble forecasting — achieving around 33% lower track error for Hurricane Helene compared to the original Aurora.

🔴 🤖 Models July 9, 2026 · 2 min read

OpenAI: GPT-5.6 arrives in three variants — Sol, Terra and Luna, with multi-agent orchestration and same-day availability in GitHub Copilot

Editorial illustration: three planetary spheres (Sol, Terra, Luna) connected by data flows

GPT-5.6 is OpenAI's new model family with three variants: Sol (flagship for complex reasoning), Terra (balanced), and Luna (high-volume, cost-efficient). It introduces Programmatic Tool Calling, explicit prompt cache control, persisted reasoning and multi-agent orchestration in beta — available from day one in GitHub Copilot.

🟡 🤖 Models July 9, 2026 · 2 min read

Meta: Muse Spark 1.1 brings multimodal reasoning with one million token context and sub-agent coordination via Model API

Editorial illustration: a central spark branching into multiple smaller agents and application flows

Muse Spark 1.1 is a multimodal reasoning model from Meta Superintelligence Labs for agentic tasks, with a context window of one million tokens. The model coordinates sub-agents, navigates flows across multiple apps, and accepts images, video, and PDFs — available through the Meta Model API in public preview and in Meta AI's Thinking mode.

🏥 In Practice

More in In Practice
🟡 🏥 In Practice July 10, 2026 · 2 min read

AWS: SageMaker Brings Serverless Fine-Tuning for NVIDIA Nemotron 3 Models with SFT, RLVR, and RLAIF Techniques

Editorial illustration: a modular model being refined through three automated pipelines with no servers in sight

Amazon SageMaker AI has introduced serverless customization for NVIDIA Nemotron 3 models, requiring no infrastructure management. Three techniques are available: SFT (supervised fine-tuning), RLVR (reinforcement learning with verifiable rewards), and RLAIF (reinforcement learning from AI feedback), making advanced RL methods accessible to enterprise teams without ML infrastructure expertise.

🟡 🏥 In Practice July 10, 2026 · 2 min read

GitHub: Better Tools Made Copilot Code Review Worse — Rewriting Prompts Restored Quality at 20% Lower Cost

Editorial illustration: a magnifying glass focused on a colored diff instead of an entire repository

GitHub revealed that migrating Copilot code review to better-maintained tools initially made results worse — the cause was not the tools but outdated agent instructions. By rewriting the instructions to a diff-first approach (batch discovery before reading files, analysis anchored to the PR diff), they achieved around 20% lower average review cost with the same quality.

🟡 🏥 In Practice July 10, 2026 · 2 min read

OpenAI: Deutsche Telekom launches AI-native transformation — partnership covers support, network and voice AI for ~245 million customers

Editorial illustration: a telecommunications tower radiating AI signals across Europe

Deutsche Telekom and OpenAI have formed a partnership to transform the operator into an AI-native telco, with applications in customer support, employee tools, network operations and voice AI. Deutsche Telekom is one of Europe's largest operators with around 245 million customers globally, making this one of the biggest AI deals in telecommunications.

🟢 🏥 In Practice July 10, 2026 · 2 min read

Anthropic: Claude Code v2.1.206 Brings /cd with Path Suggestions, /doctor Advice for CLAUDE.md, and Auto Git Push in /commit-push-pr

Editorial illustration: a terminal window with commands and an arrow pushing a branch to a remote

Claude Code v2.1.206 is a new release of Anthropic's CLI tool that introduces a /cd command with path suggestions, a /doctor check that recommends trimming CLAUDE.md files, and automatic git push approval to the configured remote in /commit-push-pr. Background agents now update immediately after an upgrade, without a slow update on the next connection.

🟢 🏥 In Practice July 10, 2026 · 2 min read

AWS: Henry Schein One Verifies Dental X-Ray Quality with Real-Time AI — 11 Million Images per Week at 1.4-Second Latency

Editorial illustration: a dental X-ray passing through an AI network with a quality badge

Henry Schein One has developed 'Image Verify', an AI system on Amazon SageMaker for real-time quality verification of dental X-ray images. The system has been deployed to more than 10,000 locations, processes over 11 million X-rays per week at an average latency of 1.4 seconds, with the goal of reducing insurance claim rejections caused by poor image quality.

⚖️ Regulation

More in Regulation

🛡️ Security

More in Security
🟡 🛡️ Security July 10, 2026 · 2 min read

arXiv:2607.08173: 'Overthinking' — Amplified Reasoning Forces Reasoning Models to Reveal Learned Secrets, New Extraction Attack at ICML 2026

Editorial illustration: a brain made of thought loops with locked documents leaking through a crack

Overthinking is a paper accepted at ICML 2026 showing that amplifying the reasoning weight in large language models can extract hidden learned information the model would not otherwise expose. The finding opens a new class of extraction attacks targeting reasoning models such as o1, DeepSeek R1, and Claude with extended thinking.

🟡 🛡️ Security July 10, 2026 · 2 min read

GitHub: CodeQL 2.26 Introduces AI Prompt Injection Detection — First Mainstream SAST Tool to Treat AI Attacks on Par with Classic Ones

Editorial illustration: a code scanner catching a malicious message embedded in an AI prompt

CodeQL 2.26.0 is the new version of GitHub's static security analysis tool, introducing detection of AI prompt injection attacks as a new analysis type, along with support for Kotlin 2.4.0. It is the first integration of an AI-specific attack vector into a mainstream SAST tool, bringing prompt injection into the same security workflows as XSS and SQL injection.

🟡 🛡️ Security July 9, 2026 · 2 min read

OpenAI: first Bio Bug Bounty — researchers invited to find biosecurity vulnerabilities in the GPT-5.5 model

Editorial illustration: shield with DNA helix surrounded by researchers' magnifying glasses

Bio Bug Bounty is OpenAI's program in which external researchers look for biological security vulnerabilities in GPT-5.5 — ways the model might provide dangerous bioscience information despite its safeguards. This is the first formal bug bounty program from a major AI lab dedicated exclusively to biosecurity, modeled on bug bounty practices from cybersecurity.

🟢 🛡️ Security July 8, 2026 · 4 min read

DT-Guard: 4B-Parameter Model Outperforms 8B Guardrails by Separating Reasoning Supervision from Inference Speed

Editorial illustration: DT-Guard fast safety framework for large language models without chain-of-thought reasoning

DT-Guard resolves the fundamental dilemma of safety guardrail systems — speed versus robustness — by training with reasoning supervision but inferring only structured safety labels. The 4B-parameter model achieves dual-side F1=0.878, outperforming 8B baseline models.

📦 Open Source

More in Open Source
🟢 📦 Open Source July 10, 2026 · 2 min read

PyTorch: 'free normalization' fuses Layer Norm into GEMM and attention kernels — Meta targets lower training cost for large models

Editorial illustration: two GPU kernels merging into a single flow with no intermediate step

The PyTorch team at Meta published techniques for fusing normalization operations directly into GEMM and attention kernels, with the goal of eliminating their computational cost — 'free normalization'. The code is available on GitHub through a kernel library, and the direct result is faster training and inference for models with frequent Layer Norm and RMS Norm operations.

🟡 📦 Open Source July 9, 2026 · 2 min read

Ollama: $88 million in funding for local AI — 8.9 million developers and Docker's founder among investors

Editorial illustration: llama silhouette made of processor chips surrounded by local computers

Ollama is a platform for running open AI models locally that announced $88 million in funding with investors including Benchmark, Theory Ventures, and Solomon Hykes, Docker's founder. The platform serves 8.9 million developers and 85% of Fortune 500 companies, while its cloud service doubles token volume every month.

🔴 📦 Open Source July 8, 2026 · 5 min read

PyTorch 2.13: Up to 4× Less GPU Memory for LLM Training and 12.3× Faster FlexAttention

Editorial illustration: PyTorch 2.13 new release with key improvements for AI training and inference

PyTorch 2.13 ships with 3,328 commits from 526 contributors. Key highlights: nn.LinearCrossEntropyLoss cuts peak GPU memory footprint by up to 4× for LLM training, FlexAttention on Apple Silicon achieves up to 12.3× speedup, and the new torchcomms backend modernizes distributed training.

🟡 📦 Open Source July 8, 2026 · 4 min read

CNCF White Paper: Data Storage Remains the Main Barrier to Cloud-Native AI at Scale

Editorial illustration: CNCF white paper on data storage for cloud-native AI and Kubernetes infrastructure

The CNCF TAG Infrastructure published a white paper mapping data storage bottlenecks in cloud-native AI environments, distinguishing the requirements of training, inference, and agentic AI phases, and providing architectural guidelines for Kubernetes AI/ML deployments.

🟢 📦 Open Source July 8, 2026 · 3 min read

IBM and Red Hat Expand Lightwell — 6,500+ Remediated Open-Source Dependencies and Clearinghouse for Finance

Editorial illustration: IBM and Red Hat Lightwell AI system for vulnerability remediation in open-source code

IBM and Red Hat launched Lightwell Network in general availability with a catalog of 6,500+ digitally signed remediated dependencies in Java and Python, plus Lightwell Clearinghouse Premier for coordinating patch embargos in financial services, backed by a $5 billion commitment to open-source security.