<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>24 AI</title><description>24 AI delivers daily artificial-intelligence news from primary sources — models, agents, regulation, security, hardware and open source, in six languages.</description><link>https://24-ai.news/</link><language>en</language><atom:link href="https://24-ai.news/en/rss.xml" rel="self" type="application/rss+xml"/><lastBuildDate>Fri, 10 Jul 2026 23:15:50 GMT</lastBuildDate><generator>24 AI Pipeline</generator><item><title>AMD: SGLang Diffusion on ROCm Brings Image Generation and Editing to Instinct GPUs — LLM Inference Framework Expands to Diffusion</title><link>https://24-ai.news/en/news/2026-07-10/amd-sglang-diffusion-rocm/</link><guid isPermaLink="true">https://24-ai.news/en/news/2026-07-10/amd-sglang-diffusion-rocm/</guid><description>AMD has published a guide for running diffusion models for image generation and editing on Instinct GPUs via SGLang Diffusion on the ROCm stack. SGLang, an inference framework originally popular for large language models, now extends support to image diffusion, strengthening AMD&apos;s AI inference offering beyond the NVIDIA ecosystem.</description><pubDate>Fri, 10 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;AMD has published a guide for running diffusion models for image generation and editing on Instinct GPUs via SGLang Diffusion on the ROCm stack. SGLang, an inference framework originally popular for large language models, now extends support to image diffusion, strengthening AMD&apos;s AI inference offering beyond the NVIDIA ecosystem.&lt;/strong&gt;&lt;/p&gt;</content:encoded><category>hardware</category><category>zanimljivo</category></item><item><title>arXiv:2607.08403: Game Theory Against Hallucinations — Multi-Agent Framework Models LLMs as Strategic Players to Reduce Confabulation</title><link>https://24-ai.news/en/news/2026-07-10/arxiv-game-theory-llm-hallucination/</link><guid isPermaLink="true">https://24-ai.news/en/news/2026-07-10/arxiv-game-theory-llm-hallucination/</guid><description>A new paper proposes a multi-agent framework using game theory to reduce hallucinations in large language models. Model interactions are modeled as a game of strategic players where equilibrium corresponds to consistent, non-hallucinated responses — offering a path distinct from standard RLHF and RAG approaches.</description><pubDate>Fri, 10 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;A new paper proposes a multi-agent framework using game theory to reduce hallucinations in large language models. Model interactions are modeled as a game of strategic players where equilibrium corresponds to consistent, non-hallucinated responses — offering a path distinct from standard RLHF and RAG approaches.&lt;/strong&gt;&lt;/p&gt;</content:encoded><category>agents</category><category>zanimljivo</category></item><item><title>arXiv:2607.08173: &apos;Overthinking&apos; — Amplified Reasoning Forces Reasoning Models to Reveal Learned Secrets, New Extraction Attack at ICML 2026</title><link>https://24-ai.news/en/news/2026-07-10/arxiv-overthinking-reasoning-secrets/</link><guid isPermaLink="true">https://24-ai.news/en/news/2026-07-10/arxiv-overthinking-reasoning-secrets/</guid><description>Overthinking is a paper accepted at ICML 2026 showing that amplifying the reasoning weight in large language models can extract hidden learned information the model would not otherwise expose. The finding opens a new class of extraction attacks targeting reasoning models such as o1, DeepSeek R1, and Claude with extended thinking.</description><pubDate>Fri, 10 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;Overthinking is a paper accepted at ICML 2026 showing that amplifying the reasoning weight in large language models can extract hidden learned information the model would not otherwise expose. The finding opens a new class of extraction attacks targeting reasoning models such as o1, DeepSeek R1, and Claude with extended thinking.&lt;/strong&gt;&lt;/p&gt;</content:encoded><category>security</category><category>važno</category></item><item><title>arXiv:2607.08733: &apos;Super Weights&apos; Explain Why Selective Fine-Tuning Fails — Paper Accepted at COLM 2026</title><link>https://24-ai.news/en/news/2026-07-10/arxiv-super-weights-selective-training/</link><guid isPermaLink="true">https://24-ai.news/en/news/2026-07-10/arxiv-super-weights-selective-training/</guid><description>Super Weights in LLMs is a paper accepted at COLM 2026 that identifies &apos;super weights&apos; — a small number of parameters whose modification disproportionately changes a language model&apos;s behavior. The paper shows that these weights explain why selective fine-tuning of only certain layers often fails, with direct implications for PEFT and LoRA approaches.</description><pubDate>Fri, 10 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;Super Weights in LLMs is a paper accepted at COLM 2026 that identifies &apos;super weights&apos; — a small number of parameters whose modification disproportionately changes a language model&apos;s behavior. The paper shows that these weights explain why selective fine-tuning of only certain layers often fails, with direct implications for PEFT and LoRA approaches.&lt;/strong&gt;&lt;/p&gt;</content:encoded><category>models</category><category>važno</category></item><item><title>AWS: SageMaker Brings Serverless Fine-Tuning for NVIDIA Nemotron 3 Models with SFT, RLVR, and RLAIF Techniques</title><link>https://24-ai.news/en/news/2026-07-10/aws-sagemaker-serverless-nemotron-finetune/</link><guid isPermaLink="true">https://24-ai.news/en/news/2026-07-10/aws-sagemaker-serverless-nemotron-finetune/</guid><description>Amazon SageMaker AI has introduced serverless customization for NVIDIA Nemotron 3 models, requiring no infrastructure management. Three techniques are available: SFT (supervised fine-tuning), RLVR (reinforcement learning with verifiable rewards), and RLAIF (reinforcement learning from AI feedback), making advanced RL methods accessible to enterprise teams without ML infrastructure expertise.</description><pubDate>Fri, 10 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;Amazon SageMaker AI has introduced serverless customization for NVIDIA Nemotron 3 models, requiring no infrastructure management. Three techniques are available: SFT (supervised fine-tuning), RLVR (reinforcement learning with verifiable rewards), and RLAIF (reinforcement learning from AI feedback), making advanced RL methods accessible to enterprise teams without ML infrastructure expertise.&lt;/strong&gt;&lt;/p&gt;</content:encoded><category>practice</category><category>važno</category></item><item><title>Anthropic: Claude Code v2.1.206 Brings /cd with Path Suggestions, /doctor Advice for CLAUDE.md, and Auto Git Push in /commit-push-pr</title><link>https://24-ai.news/en/news/2026-07-10/claude-code-2-1-206-release/</link><guid isPermaLink="true">https://24-ai.news/en/news/2026-07-10/claude-code-2-1-206-release/</guid><description>Claude Code v2.1.206 is a new release of Anthropic&apos;s CLI tool that introduces a /cd command with path suggestions, a /doctor check that recommends trimming CLAUDE.md files, and automatic git push approval to the configured remote in /commit-push-pr. Background agents now update immediately after an upgrade, without a slow update on the next connection.</description><pubDate>Fri, 10 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;Claude Code v2.1.206 is a new release of Anthropic&apos;s CLI tool that introduces a /cd command with path suggestions, a /doctor check that recommends trimming CLAUDE.md files, and automatic git push approval to the configured remote in /commit-push-pr. Background agents now update immediately after an upgrade, without a slow update on the next connection.&lt;/strong&gt;&lt;/p&gt;</content:encoded><category>practice</category><category>zanimljivo</category></item><item><title>GitHub: CodeQL 2.26 Introduces AI Prompt Injection Detection — First Mainstream SAST Tool to Treat AI Attacks on Par with Classic Ones</title><link>https://24-ai.news/en/news/2026-07-10/github-codeql-prompt-injection-detection/</link><guid isPermaLink="true">https://24-ai.news/en/news/2026-07-10/github-codeql-prompt-injection-detection/</guid><description>CodeQL 2.26.0 is the new version of GitHub&apos;s static security analysis tool, introducing detection of AI prompt injection attacks as a new analysis type, along with support for Kotlin 2.4.0. It is the first integration of an AI-specific attack vector into a mainstream SAST tool, bringing prompt injection into the same security workflows as XSS and SQL injection.</description><pubDate>Fri, 10 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;CodeQL 2.26.0 is the new version of GitHub&apos;s static security analysis tool, introducing detection of AI prompt injection attacks as a new analysis type, along with support for Kotlin 2.4.0. It is the first integration of an AI-specific attack vector into a mainstream SAST tool, bringing prompt injection into the same security workflows as XSS and SQL injection.&lt;/strong&gt;&lt;/p&gt;</content:encoded><category>security</category><category>važno</category></item><item><title>GitHub: Better Tools Made Copilot Code Review Worse — Rewriting Prompts Restored Quality at 20% Lower Cost</title><link>https://24-ai.news/en/news/2026-07-10/github-copilot-code-review-prompt-fix/</link><guid isPermaLink="true">https://24-ai.news/en/news/2026-07-10/github-copilot-code-review-prompt-fix/</guid><description>GitHub revealed that migrating Copilot code review to better-maintained tools initially made results worse — the cause was not the tools but outdated agent instructions. By rewriting the instructions to a diff-first approach (batch discovery before reading files, analysis anchored to the PR diff), they achieved around 20% lower average review cost with the same quality.</description><pubDate>Fri, 10 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;GitHub revealed that migrating Copilot code review to better-maintained tools initially made results worse — the cause was not the tools but outdated agent instructions. By rewriting the instructions to a diff-first approach (batch discovery before reading files, analysis anchored to the PR diff), they achieved around 20% lower average review cost with the same quality.&lt;/strong&gt;&lt;/p&gt;</content:encoded><category>practice</category><category>važno</category></item><item><title>AWS: Henry Schein One Verifies Dental X-Ray Quality with Real-Time AI — 11 Million Images per Week at 1.4-Second Latency</title><link>https://24-ai.news/en/news/2026-07-10/henry-schein-dental-ai-image-verify/</link><guid isPermaLink="true">https://24-ai.news/en/news/2026-07-10/henry-schein-dental-ai-image-verify/</guid><description>Henry Schein One has developed &apos;Image Verify&apos;, an AI system on Amazon SageMaker for real-time quality verification of dental X-ray images. The system has been deployed to more than 10,000 locations, processes over 11 million X-rays per week at an average latency of 1.4 seconds, with the goal of reducing insurance claim rejections caused by poor image quality.</description><pubDate>Fri, 10 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;Henry Schein One has developed &apos;Image Verify&apos;, an AI system on Amazon SageMaker for real-time quality verification of dental X-ray images. The system has been deployed to more than 10,000 locations, processes over 11 million X-rays per week at an average latency of 1.4 seconds, with the goal of reducing insurance claim rejections caused by poor image quality.&lt;/strong&gt;&lt;/p&gt;</content:encoded><category>practice</category><category>zanimljivo</category></item><item><title>LangChain: OpenWiki Brains gives agents proactive memory — automatically collects context from Gmail, Notion and Git into Markdown</title><link>https://24-ai.news/en/news/2026-07-10/langchain-openwiki-brains-agent-memory/</link><guid isPermaLink="true">https://24-ai.news/en/news/2026-07-10/langchain-openwiki-brains-agent-memory/</guid><description>OpenWiki Brains is LangChain&apos;s open-source framework that gives AI agents proactive memory: it automatically collects context from external sources and saves it as Markdown files on disk. The first version supports 6 connectors — Gmail, Notion, Git, Twitter/X, Hacker News and web search — and runs locally through scheduled jobs instead of relying on the cloud.</description><pubDate>Fri, 10 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;OpenWiki Brains is LangChain&apos;s open-source framework that gives AI agents proactive memory: it automatically collects context from external sources and saves it as Markdown files on disk. The first version supports 6 connectors — Gmail, Notion, Git, Twitter/X, Hacker News and web search — and runs locally through scheduled jobs instead of relying on the cloud.&lt;/strong&gt;&lt;/p&gt;</content:encoded><category>agents</category><category>važno</category></item><item><title>OpenAI: Deutsche Telekom launches AI-native transformation — partnership covers support, network and voice AI for ~245 million customers</title><link>https://24-ai.news/en/news/2026-07-10/openai-deutsche-telekom-ai-native-telco/</link><guid isPermaLink="true">https://24-ai.news/en/news/2026-07-10/openai-deutsche-telekom-ai-native-telco/</guid><description>Deutsche Telekom and OpenAI have formed a partnership to transform the operator into an AI-native telco, with applications in customer support, employee tools, network operations and voice AI. Deutsche Telekom is one of Europe&apos;s largest operators with around 245 million customers globally, making this one of the biggest AI deals in telecommunications.</description><pubDate>Fri, 10 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;Deutsche Telekom and OpenAI have formed a partnership to transform the operator into an AI-native telco, with applications in customer support, employee tools, network operations and voice AI. Deutsche Telekom is one of Europe&apos;s largest operators with around 245 million customers globally, making this one of the biggest AI deals in telecommunications.&lt;/strong&gt;&lt;/p&gt;</content:encoded><category>practice</category><category>važno</category></item><item><title>PyTorch: &apos;free normalization&apos; fuses Layer Norm into GEMM and attention kernels — Meta targets lower training cost for large models</title><link>https://24-ai.news/en/news/2026-07-10/pytorch-free-normalization-kernel-fusion/</link><guid isPermaLink="true">https://24-ai.news/en/news/2026-07-10/pytorch-free-normalization-kernel-fusion/</guid><description>The PyTorch team at Meta published techniques for fusing normalization operations directly into GEMM and attention kernels, with the goal of eliminating their computational cost — &apos;free normalization&apos;. The code is available on GitHub through a kernel library, and the direct result is faster training and inference for models with frequent Layer Norm and RMS Norm operations.</description><pubDate>Fri, 10 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;The PyTorch team at Meta published techniques for fusing normalization operations directly into GEMM and attention kernels, with the goal of eliminating their computational cost — &apos;free normalization&apos;. The code is available on GitHub through a kernel library, and the direct result is faster training and inference for models with frequent Layer Norm and RMS Norm operations.&lt;/strong&gt;&lt;/p&gt;</content:encoded><category>open-source</category><category>zanimljivo</category></item><item><title>AMD: FlyDSL — Python DSL that promises performance of hand-written HIP C++ GPU kernels with far less code</title><link>https://24-ai.news/en/news/2026-07-09/amd-flydsl-python-gpu-kernels/</link><guid isPermaLink="true">https://24-ai.news/en/news/2026-07-09/amd-flydsl-python-gpu-kernels/</guid><description>FlyDSL is AMD&apos;s Python domain-specific language for writing GPU kernels that, according to a new ROCm guide, matches the performance of manually optimized HIP C++ code while requiring fewer written lines. AMD published a practical guide for porting existing HIP kernels, positioning FlyDSL as a Python-first alternative to NVIDIA&apos;s CUDA toolchain.</description><pubDate>Thu, 09 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;FlyDSL is AMD&apos;s Python domain-specific language for writing GPU kernels that, according to a new ROCm guide, matches the performance of manually optimized HIP C++ code while requiring fewer written lines. AMD published a practical guide for porting existing HIP kernels, positioning FlyDSL as a Python-first alternative to NVIDIA&apos;s CUDA toolchain.&lt;/strong&gt;&lt;/p&gt;</content:encoded><category>hardware</category><category>zanimljivo</category></item><item><title>Anthropic: Nobel laureate Ben Bernanke, former Fed chair, joins the Long-Term Benefit Trust — the body overseeing the company&apos;s mission</title><link>https://24-ai.news/en/news/2026-07-09/anthropic-bernanke-benefit-trust/</link><guid isPermaLink="true">https://24-ai.news/en/news/2026-07-09/anthropic-bernanke-benefit-trust/</guid><description>Ben Bernanke is the former chair of the US Federal Reserve (2006–2014) and 2022 Nobel laureate in economics, appointed by Anthropic to its Long-Term Benefit Trust. The LTBT is an independent body with no equity or profit interest that oversees the company&apos;s public benefit mission — Bernanke brings expertise on the economic impacts of AI.</description><pubDate>Thu, 09 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;Ben Bernanke is the former chair of the US Federal Reserve (2006–2014) and 2022 Nobel laureate in economics, appointed by Anthropic to its Long-Term Benefit Trust. The LTBT is an independent body with no equity or profit interest that oversees the company&apos;s public benefit mission — Bernanke brings expertise on the economic impacts of AI.&lt;/strong&gt;&lt;/p&gt;</content:encoded><category>community</category><category>važno</category></item><item><title>Anthropic: &apos;Hard Questions&apos; initiative invites the public to ask the toughest questions about AI — with a commitment to public reporting</title><link>https://24-ai.news/en/news/2026-07-09/anthropic-hard-questions-initiative/</link><guid isPermaLink="true">https://24-ai.news/en/news/2026-07-09/anthropic-hard-questions-initiative/</guid><description>Hard Questions is Anthropic&apos;s initiative through which the public can ask the toughest questions about artificial intelligence, and the company commits to transparently tracking and publicly reporting on concrete actions taken. It is based on a survey of 52,000 Americans and the views of 81,000 Claude users from 159 countries in 70 languages.</description><pubDate>Thu, 09 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;Hard Questions is Anthropic&apos;s initiative through which the public can ask the toughest questions about artificial intelligence, and the company commits to transparently tracking and publicly reporting on concrete actions taken. It is based on a survey of 52,000 Americans and the views of 81,000 Claude users from 159 countries in 70 languages.&lt;/strong&gt;&lt;/p&gt;</content:encoded><category>community</category><category>zanimljivo</category></item><item><title>Anthropic: Reflect — dashboard showing how you use Claude, with quiet hours and self-reflection prompts</title><link>https://24-ai.news/en/news/2026-07-09/anthropic-reflect-usage-dashboard/</link><guid isPermaLink="true">https://24-ai.news/en/news/2026-07-09/anthropic-reflect-usage-dashboard/</guid><description>Reflect is Anthropic&apos;s beta dashboard for analyzing your own Claude usage habits: it shows dominant topics, task categories, and time patterns over periods of 1 to 12 months. Users can set quiet hours and break reminders, and it&apos;s available on Free, Pro, and Max plans with Memory enabled.</description><pubDate>Thu, 09 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;Reflect is Anthropic&apos;s beta dashboard for analyzing your own Claude usage habits: it shows dominant topics, task categories, and time patterns over periods of 1 to 12 months. Users can set quiet hours and break reminders, and it&apos;s available on Free, Pro, and Max plans with Memory enabled.&lt;/strong&gt;&lt;/p&gt;</content:encoded><category>practice</category><category>zanimljivo</category></item><item><title>arXiv:2607.06906: &apos;Harness Effect&apos; — orchestration design cuts AI agent cost by 41%, more than changing the model</title><link>https://24-ai.news/en/news/2026-07-09/arxiv-harness-effect-token-economics/</link><guid isPermaLink="true">https://24-ai.news/en/news/2026-07-09/arxiv-harness-effect-token-economics/</guid><description>The Harness Effect is a finding from an empirical paper by 32 authors showing that the design of the orchestration layer affects AI agent costs more than the choice of the model itself. Across 22 tasks and 6 models, optimized orchestration reduced cost per task by 41% (from $0.21 to $0.12), tokens by 38%, and execution time by 44% — while also improving quality.</description><pubDate>Thu, 09 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;The Harness Effect is a finding from an empirical paper by 32 authors showing that the design of the orchestration layer affects AI agent costs more than the choice of the model itself. Across 22 tasks and 6 models, optimized orchestration reduced cost per task by 41% (from $0.21 to $0.12), tokens by 38%, and execution time by 44% — while also improving quality.&lt;/strong&gt;&lt;/p&gt;</content:encoded><category>agents</category><category>važno</category></item><item><title>GitHub: Copilot now offers AI repository overview — automatic summary of purpose, technologies, and how to contribute</title><link>https://24-ai.news/en/news/2026-07-09/github-copilot-repository-overview/</link><guid isPermaLink="true">https://24-ai.news/en/news/2026-07-09/github-copilot-repository-overview/</guid><description>GitHub Copilot repository overview is a new feature that automatically generates a summary when first visiting an unfamiliar repository: what the project does, what technologies it uses, and how to contribute. It works on github.com through Copilot chat, is available on all Copilot plans, and requires no configuration.</description><pubDate>Thu, 09 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;GitHub Copilot repository overview is a new feature that automatically generates a summary when first visiting an unfamiliar repository: what the project does, what technologies it uses, and how to contribute. It works on github.com through Copilot chat, is available on all Copilot plans, and requires no configuration.&lt;/strong&gt;&lt;/p&gt;</content:encoded><category>practice</category><category>važno</category></item><item><title>Google: SensorFM — foundation model trained on one trillion minutes of wearable data wins on 34 of 35 health tasks</title><link>https://24-ai.news/en/news/2026-07-09/google-sensorfm-wearable-health/</link><guid isPermaLink="true">https://24-ai.news/en/news/2026-07-09/google-sensorfm-wearable-health/</guid><description>SensorFM is Google&apos;s foundation model for health data from wearable devices, trained on more than one trillion minutes of signals from Fitbit and Pixel Watch devices worn by 5 million users in over 100 countries. The model outperforms specialized approaches on 34 of 35 tasks, with +9% AUC on classification and +21% correlation on regression.</description><pubDate>Thu, 09 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;SensorFM is Google&apos;s foundation model for health data from wearable devices, trained on more than one trillion minutes of signals from Fitbit and Pixel Watch devices worn by 5 million users in over 100 countries. The model outperforms specialized approaches on 34 of 35 tasks, with +9% AUC on classification and +21% correlation on regression.&lt;/strong&gt;&lt;/p&gt;</content:encoded><category>models</category><category>kritično</category></item><item><title>IBM: Bob becomes a multi-agent platform — nine-month COBOL modernization project completed in 3 days</title><link>https://24-ai.news/en/news/2026-07-09/ibm-bob-multi-agent-mainframe/</link><guid isPermaLink="true">https://24-ai.news/en/news/2026-07-09/ibm-bob-multi-agent-mainframe/</guid><description>IBM Bob is an agentic platform for software development that in its new version gains a multi-agent architecture for the full development lifecycle, Bobalytics analytics, and a Premium package for mainframe modernization of COBOL and PL/I. IBM reports that client Blue Pearl completed a nine-month modernization project with Bob in just three days.</description><pubDate>Thu, 09 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;IBM Bob is an agentic platform for software development that in its new version gains a multi-agent architecture for the full development lifecycle, Bobalytics analytics, and a Premium package for mainframe modernization of COBOL and PL/I. IBM reports that client Blue Pearl completed a nine-month modernization project with Bob in just three days.&lt;/strong&gt;&lt;/p&gt;</content:encoded><category>agents</category><category>važno</category></item><item><title>Meta: Muse Spark 1.1 brings multimodal reasoning with one million token context and sub-agent coordination via Model API</title><link>https://24-ai.news/en/news/2026-07-09/meta-muse-spark-1-1-reasoning/</link><guid isPermaLink="true">https://24-ai.news/en/news/2026-07-09/meta-muse-spark-1-1-reasoning/</guid><description>Muse Spark 1.1 is a multimodal reasoning model from Meta Superintelligence Labs for agentic tasks, with a context window of one million tokens. The model coordinates sub-agents, navigates flows across multiple apps, and accepts images, video, and PDFs — available through the Meta Model API in public preview and in Meta AI&apos;s Thinking mode.</description><pubDate>Thu, 09 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;Muse Spark 1.1 is a multimodal reasoning model from Meta Superintelligence Labs for agentic tasks, with a context window of one million tokens. The model coordinates sub-agents, navigates flows across multiple apps, and accepts images, video, and PDFs — available through the Meta Model API in public preview and in Meta AI&apos;s Thinking mode.&lt;/strong&gt;&lt;/p&gt;</content:encoded><category>models</category><category>važno</category></item><item><title>Microsoft: Aurora 1.5 open-source model outperforms ECMWF ensemble on 88.9% of variables — new standard in AI weather forecasting</title><link>https://24-ai.news/en/news/2026-07-09/microsoft-aurora-1-5-weather-sota/</link><guid isPermaLink="true">https://24-ai.news/en/news/2026-07-09/microsoft-aurora-1-5-weather-sota/</guid><description>Aurora 1.5 is Microsoft&apos;s open-source foundation model for the Earth system that outperforms the ECMWF ensemble forecast on 88.9% of evaluated variables and horizons. The new release adds 22 meteorological variables, hourly resolution, and probabilistic ensemble forecasting — achieving around 33% lower track error for Hurricane Helene compared to the original Aurora.</description><pubDate>Thu, 09 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;Aurora 1.5 is Microsoft&apos;s open-source foundation model for the Earth system that outperforms the ECMWF ensemble forecast on 88.9% of evaluated variables and horizons. The new release adds 22 meteorological variables, hourly resolution, and probabilistic ensemble forecasting — achieving around 33% lower track error for Hurricane Helene compared to the original Aurora.&lt;/strong&gt;&lt;/p&gt;</content:encoded><category>models</category><category>kritično</category></item><item><title>Mistral: Studio gets version control for prompts and skills — immutable versions, audit trails, and rollback</title><link>https://24-ai.news/en/news/2026-07-09/mistral-studio-prompt-skill-governance/</link><guid isPermaLink="true">https://24-ai.news/en/news/2026-07-09/mistral-studio-prompt-skill-governance/</guid><description>Mistral Studio introduces a system of record for AI prompts and skills: immutable versioning with full history, ownership and audit trails, classification labels, and quick rollback. Skills are integrated as MCP servers, and development and production flows are separated — changes in production go through CI/CD.</description><pubDate>Thu, 09 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;Mistral Studio introduces a system of record for AI prompts and skills: immutable versioning with full history, ownership and audit trails, classification labels, and quick rollback. Skills are integrated as MCP servers, and development and production flows are separated — changes in production go through CI/CD.&lt;/strong&gt;&lt;/p&gt;</content:encoded><category>practice</category><category>zanimljivo</category></item><item><title>Ollama: $88 million in funding for local AI — 8.9 million developers and Docker&apos;s founder among investors</title><link>https://24-ai.news/en/news/2026-07-09/ollama-88m-funding-local-ai/</link><guid isPermaLink="true">https://24-ai.news/en/news/2026-07-09/ollama-88m-funding-local-ai/</guid><description>Ollama is a platform for running open AI models locally that announced $88 million in funding with investors including Benchmark, Theory Ventures, and Solomon Hykes, Docker&apos;s founder. The platform serves 8.9 million developers and 85% of Fortune 500 companies, while its cloud service doubles token volume every month.</description><pubDate>Thu, 09 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;Ollama is a platform for running open AI models locally that announced $88 million in funding with investors including Benchmark, Theory Ventures, and Solomon Hykes, Docker&apos;s founder. The platform serves 8.9 million developers and 85% of Fortune 500 companies, while its cloud service doubles token volume every month.&lt;/strong&gt;&lt;/p&gt;</content:encoded><category>open-source</category><category>važno</category></item><item><title>OpenAI: first Bio Bug Bounty — researchers invited to find biosecurity vulnerabilities in the GPT-5.5 model</title><link>https://24-ai.news/en/news/2026-07-09/openai-bio-bug-bounty-gpt-5-5/</link><guid isPermaLink="true">https://24-ai.news/en/news/2026-07-09/openai-bio-bug-bounty-gpt-5-5/</guid><description>Bio Bug Bounty is OpenAI&apos;s program in which external researchers look for biological security vulnerabilities in GPT-5.5 — ways the model might provide dangerous bioscience information despite its safeguards. This is the first formal bug bounty program from a major AI lab dedicated exclusively to biosecurity, modeled on bug bounty practices from cybersecurity.</description><pubDate>Thu, 09 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;Bio Bug Bounty is OpenAI&apos;s program in which external researchers look for biological security vulnerabilities in GPT-5.5 — ways the model might provide dangerous bioscience information despite its safeguards. This is the first formal bug bounty program from a major AI lab dedicated exclusively to biosecurity, modeled on bug bounty practices from cybersecurity.&lt;/strong&gt;&lt;/p&gt;</content:encoded><category>security</category><category>važno</category></item><item><title>OpenAI: ChatGPT becomes an agent for serious work — operates for hours through apps and files until the job is done</title><link>https://24-ai.news/en/news/2026-07-09/openai-chatgpt-work-agent/</link><guid isPermaLink="true">https://24-ai.news/en/news/2026-07-09/openai-chatgpt-work-agent/</guid><description>ChatGPT Work is OpenAI&apos;s move into agentic work: ChatGPT can now take actions across apps and files and work independently on a project for hours, turning a stated goal into completed work. OpenAI positions it as a partner for the most ambitious tasks, competing directly with agentic products from Anthropic and GitHub.</description><pubDate>Thu, 09 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;ChatGPT Work is OpenAI&apos;s move into agentic work: ChatGPT can now take actions across apps and files and work independently on a project for hours, turning a stated goal into completed work. OpenAI positions it as a partner for the most ambitious tasks, competing directly with agentic products from Anthropic and GitHub.&lt;/strong&gt;&lt;/p&gt;</content:encoded><category>agents</category><category>važno</category></item><item><title>OpenAI: GPT-5.6 arrives in three variants — Sol, Terra and Luna, with multi-agent orchestration and same-day availability in GitHub Copilot</title><link>https://24-ai.news/en/news/2026-07-09/openai-gpt-5-6-sol-terra-luna/</link><guid isPermaLink="true">https://24-ai.news/en/news/2026-07-09/openai-gpt-5-6-sol-terra-luna/</guid><description>GPT-5.6 is OpenAI&apos;s new model family with three variants: Sol (flagship for complex reasoning), Terra (balanced), and Luna (high-volume, cost-efficient). It introduces Programmatic Tool Calling, explicit prompt cache control, persisted reasoning and multi-agent orchestration in beta — available from day one in GitHub Copilot.</description><pubDate>Thu, 09 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;GPT-5.6 is OpenAI&apos;s new model family with three variants: Sol (flagship for complex reasoning), Terra (balanced), and Luna (high-volume, cost-efficient). It introduces Programmatic Tool Calling, explicit prompt cache control, persisted reasoning and multi-agent orchestration in beta — available from day one in GitHub Copilot.&lt;/strong&gt;&lt;/p&gt;</content:encoded><category>models</category><category>kritično</category></item><item><title>xAI: Grok 4.5 — new flagship for coding and agentic work at $2 per million input tokens</title><link>https://24-ai.news/en/news/2026-07-09/xai-grok-4-5-coding-agentic/</link><guid isPermaLink="true">https://24-ai.news/en/news/2026-07-09/xai-grok-4-5-coding-agentic/</guid><description>Grok 4.5 is xAI&apos;s new model positioned as a flagship for coding and agentic tasks, priced at $2 per million input and $6 per million output tokens. It supports configurable reasoning effort (low/medium/high) and is already available through Perplexity&apos;s Agent API, aggressively undercutting comparable frontier models.</description><pubDate>Thu, 09 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;Grok 4.5 is xAI&apos;s new model positioned as a flagship for coding and agentic tasks, priced at $2 per million input and $6 per million output tokens. It supports configurable reasoning effort (low/medium/high) and is already available through Perplexity&apos;s Agent API, aggressively undercutting comparable frontier models.&lt;/strong&gt;&lt;/p&gt;</content:encoded><category>models</category><category>važno</category></item><item><title>Claude Code v2.1.205: Session Security, Bug Fixes, and 400 MB RAM Savings</title><link>https://24-ai.news/en/news/2026-07-08/anthropic-claude-code-2-1-205/</link><guid isPermaLink="true">https://24-ai.news/en/news/2026-07-08/anthropic-claude-code-2-1-205/</guid><description>Anthropic released Claude Code v2.1.205 on July 8, 2026 with a new security rule that blocks manipulation of session transcript files, a series of fixes including silent JSON schema and Windows worktree bugs, and an auto-update optimization that now streams the binary to disk instead of into memory — saving approximately 400 MB of peak RAM.</description><pubDate>Wed, 08 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;Anthropic released Claude Code v2.1.205 on July 8, 2026 with a new security rule that blocks manipulation of session transcript files, a series of fixes including silent JSON schema and Windows worktree bugs, and an auto-update optimization that now streams the binary to disk instead of into memory — saving approximately 400 MB of peak RAM.&lt;/strong&gt;&lt;/p&gt;</content:encoded><category>practice</category><category>zanimljivo</category></item><item><title>Anthropic Develops an &apos;Off Switch&apos; for Dangerous Knowledge: GRAM Isolates Dual-Use Capabilities into Removable Modules</title><link>https://24-ai.news/en/news/2026-07-08/anthropic-gram-dual-use-offswitch/</link><guid isPermaLink="true">https://24-ai.news/en/news/2026-07-08/anthropic-gram-dual-use-offswitch/</guid><description>Anthropic and AE Studio announce GRAM (Gradient-Routed Auxiliary Modules) — a method that isolates dual-use knowledge such as virology, cybersecurity, and nuclear physics into removable neural modules during training, enabling a single training run to produce multiple model variants with different capability sets.</description><pubDate>Wed, 08 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;Anthropic and AE Studio announce GRAM (Gradient-Routed Auxiliary Modules) — a method that isolates dual-use knowledge such as virology, cybersecurity, and nuclear physics into removable neural modules during training, enabling a single training run to produce multiple model variants with different capability sets.&lt;/strong&gt;&lt;/p&gt;</content:encoded><category>models</category><category>važno</category></item></channel></rss>