🤖 Models

189 articles

🟡 🤖 Models July 10, 2026 · 2 min read

arXiv:2607.08733: 'Super Weights' Explain Why Selective Fine-Tuning Fails — Paper Accepted at COLM 2026

Editorial illustration: a network of parameters in which a few bright nodes hold the entire structure together

Super Weights in LLMs is a paper accepted at COLM 2026 that identifies 'super weights' — a small number of parameters whose modification disproportionately changes a language model's behavior. The paper shows that these weights explain why selective fine-tuning of only certain layers often fails, with direct implications for PEFT and LoRA approaches.

🔴 🤖 Models July 9, 2026 · 2 min read

Google: SensorFM — foundation model trained on one trillion minutes of wearable data wins on 34 of 35 health tasks

Editorial illustration: smartwatch with waves of biometric signals flowing into a neural network

SensorFM is Google's foundation model for health data from wearable devices, trained on more than one trillion minutes of signals from Fitbit and Pixel Watch devices worn by 5 million users in over 100 countries. The model outperforms specialized approaches on 34 of 35 tasks, with +9% AUC on classification and +21% correlation on regression.

🔴 🤖 Models July 9, 2026 · 2 min read

Microsoft: Aurora 1.5 open-source model outperforms ECMWF ensemble on 88.9% of variables — new standard in AI weather forecasting

Editorial illustration: globe with swirling weather fronts and data network

Aurora 1.5 is Microsoft's open-source foundation model for the Earth system that outperforms the ECMWF ensemble forecast on 88.9% of evaluated variables and horizons. The new release adds 22 meteorological variables, hourly resolution, and probabilistic ensemble forecasting — achieving around 33% lower track error for Hurricane Helene compared to the original Aurora.

🔴 🤖 Models July 9, 2026 · 2 min read

OpenAI: GPT-5.6 arrives in three variants — Sol, Terra and Luna, with multi-agent orchestration and same-day availability in GitHub Copilot

Editorial illustration: three planetary spheres (Sol, Terra, Luna) connected by data flows

GPT-5.6 is OpenAI's new model family with three variants: Sol (flagship for complex reasoning), Terra (balanced), and Luna (high-volume, cost-efficient). It introduces Programmatic Tool Calling, explicit prompt cache control, persisted reasoning and multi-agent orchestration in beta — available from day one in GitHub Copilot.

🟡 🤖 Models July 9, 2026 · 2 min read

Meta: Muse Spark 1.1 brings multimodal reasoning with one million token context and sub-agent coordination via Model API

Editorial illustration: a central spark branching into multiple smaller agents and application flows

Muse Spark 1.1 is a multimodal reasoning model from Meta Superintelligence Labs for agentic tasks, with a context window of one million tokens. The model coordinates sub-agents, navigates flows across multiple apps, and accepts images, video, and PDFs — available through the Meta Model API in public preview and in Meta AI's Thinking mode.

🟡 🤖 Models July 9, 2026 · 2 min read

xAI: Grok 4.5 — new flagship for coding and agentic work at $2 per million input tokens

Editorial illustration: dark console with code from which a robotic arm with a tool emerges

Grok 4.5 is xAI's new model positioned as a flagship for coding and agentic tasks, priced at $2 per million input and $6 per million output tokens. It supports configurable reasoning effort (low/medium/high) and is already available through Perplexity's Agent API, aggressively undercutting comparable frontier models.

🔴 🤖 Models July 8, 2026 · 5 min read

Mistral Robostral Navigate: Robotic AI That Navigates Using Only an RGB Camera

Editorial illustration: Mistral Robostral embodied robot navigation model based purely on RGB vision

Mistral introduced Robostral Navigate, its first model for embodied robotic navigation with 8 billion parameters. Using only a single RGB camera — no LiDAR or depth sensors — it achieves 76.6% success on the R2R-CE benchmark for unseen environments and surpasses multi-sensor competitors by 4.5 percentage points.

🔴 🤖 Models July 8, 2026 · 4 min read

OpenAI Launches GPT-Live: Voice Model for Lifelike AI Conversations

Editorial illustration: OpenAI GPT-Live voice model integrated into the ChatGPT Voice interface

OpenAI introduced GPT-Live, a new voice model designed for natural, lifelike conversational AI interactions. Integrated into ChatGPT Voice from day one, it arrives as part of an accelerated voice model development cycle — just two days after GPT-Realtime-2.1. Technical specifications and pricing have not yet been published.

🟡 🤖 Models July 8, 2026 · 4 min read

Anthropic Develops an 'Off Switch' for Dangerous Knowledge: GRAM Isolates Dual-Use Capabilities into Removable Modules

Editorial illustration: Anthropic GRAM interchangeable knowledge modules for controlling and removing dual-use AI capabilities

Anthropic and AE Studio announce GRAM (Gradient-Routed Auxiliary Modules) — a method that isolates dual-use knowledge such as virology, cybersecurity, and nuclear physics into removable neural modules during training, enabling a single training run to produce multiple model variants with different capability sets.

🔴 🤖 Models July 7, 2026 · 5 min read

Meta Launches Muse Image and Muse Video: Agentic AI That Self-Corrects Its Own Mistakes

Editorial illustration: Meta Muse generative AI for creating images and video content

Meta Superintelligence Labs has unveiled Muse Image and Muse Video — models that operate as agents, internally invoking code and web-search tools, ranking #2 and #3 on the Arena leaderboard, with mandatory Content Seal watermarking.

🟡 🤖 Models July 7, 2026 · 4 min read

Direct-OPD: How a Weaker Model Can Train a Stronger One Without Expensive RL

Editorial illustration: weak-to-strong RL distillation as an approach to AI superalignment

Researchers from Tsinghua University and Zhipu AI proposed Direct On-Policy Distillation (Direct-OPD) — a method that transfers reinforcement learning gains from a weaker teacher model to a stronger student without imitating the final policy, achieving a jump from 48.3% to 62.4% on AIME 2024.

🔴 🤖 Models July 6, 2026 · 5 min read

Anthropic discovers J-space: an emergent internal workspace inside Claude

Editorial illustration: Anthropic interpretability research and hidden behaviors in neural networks

Anthropic researchers have identified an emergent internal structure inside Claude called J-space, discovered using a new technique called Jacobian lens (J-lens). Inspired by the neuroscientific Global Workspace Theory, J-space acts as a silent internal reasoning space invisible in the model's output, and can reveal hidden behaviors such as data fabrication, test-scenario recognition, and implanted malicious goals.

🟡 🤖 Models July 6, 2026 · 4 min read

AWS introduces rDPO — selective unlearning for Amazon Nova models

Editorial illustration: Selective unlearning in AI models via rDPO and LoRA unlearning techniques

Amazon Web Services has published a selective unlearning technique for Nova models based on Reverse Direct Preference Optimization (rDPO). Implemented via LoRA adapters without full retraining, the technique reduces unwanted refusal rates by up to 53.74 percentage points with minimal loss of utility.

🟡 🤖 Models July 6, 2026 · 4 min read

OpenAI releases GPT-Realtime-2.1 and mini variant: improved voice recognition and noise handling

Editorial illustration: OpenAI GPT Realtime voice models for real-time agentic audio communication

On July 6, 2026, OpenAI released two new realtime voice models — GPT-Realtime-2.1 and GPT-Realtime-2.1-mini — available on the existing v1/realtime endpoint. The models bring improved alphanumeric string recognition, better handling of silence and noise, and improved interruption behavior.

🟢 🤖 Models July 6, 2026 · 4 min read

MiniMax M2, M2.1, and M2.5 now available on Amazon Bedrock

Editorial illustration: AWS Bedrock platform with new open third-party MoE models

Amazon Bedrock has expanded its offering with three open-weight MiniMax models — M2 with a one-million-token context window, M2.1 designed for deep reasoning and coding, and M2.5 with a 230-billion-parameter Mixture-of-Experts architecture designed for agentic workloads.

🟡 🤖 Models July 3, 2026 · 4 min read

ReContext Improves Utilization of 128K Context Windows Without Retraining

Editorial illustration: recursive evidence replay in a 128K token long context for a language model

Researchers from the University of Illinois developed ReContext — an inference technique that recursively replays relevant evidence from a long context window and consistently improves performance across three LLM architectures over eight benchmarks, without any retraining.

🟢 🤖 Models July 3, 2026 · 4 min read

VRRL: Reinforcement Learning Forces Visual Models to Actually Use the Image During Self-Correction

Editorial illustration: vision-language model with reinforcement learning for self-reflection over scenes

Liyan Tang, Fangcong Yin, and Greg Durrett developed VRRL — a reinforcement learning framework that uses trajectory prefix masking and experience replay to force vision-language models to ground self-reflection in real visual input, achieving significantly better performance on out-of-distribution examples.

🟡 🤖 Models July 1, 2026 · 3 min read

Fable 5 and Mythos 5 Back Online: Anthropic Restores Model Access

Editorial illustration: Anthropic restores Claude Fable 5 and Mythos 5 availability for users

On July 1, Anthropic restored access to Claude Fable 5 and Claude Mythos 5, nearly three weeks after withdrawing them due to US export controls. The return follows an improved safety classifier and a proposed industry framework for rating jailbreak severity.

🟡 🤖 Models July 1, 2026 · 4 min read

NVIDIA Nemotron and OpenAI GPT OSS Models Available in AWS GovCloud with FedRAMP High Certification

Editorial illustration: NVIDIA Nemotron and OpenAI gpt-oss models on AWS Bedrock GovCloud with strict security certifications

AWS GovCloud (US) is receiving six new models on Amazon Bedrock: OpenAI open-weight gpt-oss-120b and gpt-oss-20b, and four NVIDIA Nemotron models with a 1M token context. The infrastructure meets FedRAMP High, DoD IL 2/4/5, ITAR, and CJIS requirements with a zero operator-access design.

🟡 🤖 Models July 1, 2026 · 4 min read

GitHub Copilot Vision and Browser Tools: Two GA Capabilities in One Day

Editorial illustration: GitHub Copilot Vision and browser tools become generally available

GitHub has declared GA two Copilot capabilities: Vision for attaching images and PDFs to chat prompts, and browser tools that give agents in VS Code control over a real browser. Both are available to all plans without admin action.

🔴 🤖 Models June 30, 2026 · 4 min read

Claude Sonnet 5: Anthropic's Most Agentic Model Becomes the New Standard

Editorial illustration: Anthropic introduces the new flagship Claude Sonnet 5 model for agentic tasks

Anthropic today launches Claude Sonnet 5 — a model that plans multi-step tasks, controls browsers and terminals, and approaches Opus 4.8 performance on key benchmarks, with an introductory price of $2/$10 per million tokens through August 31, 2026, and a 1M token context window.

🟡 🤖 Models June 30, 2026 · 4 min read

Fable 5 Returns: Government Lifts Export Controls, Global Redeployment Begins July 1

Editorial illustration: Anthropic reactivates Fable 5 after the U.S. government lifts the export ban

After 18 days of global suspension, Anthropic announces the redeployment of Fable 5: the U.S. government lifted export controls on June 30, and an improved safety classifier blocks the discovered bypass technique in more than 99 percent of cases.

🟡 🤖 Models June 30, 2026 · 4 min read

Google Launches Two Models on the Same Day: Gemini Omni Flash and Nano Banana 2 Lite

Editorial illustration: Google launches Gemini Omni Flash for video and Nano Banana 2 for image generation

On June 30, 2026, Google simultaneously announced Gemini Omni Flash for conversational video editing and Nano Banana 2 Lite for rapid image synthesis. The video model is priced at $0.10 per second of output, while the image model costs just $0.034 per thousand images.

🟡 🤖 Models June 30, 2026 · 3 min read

Google Research Introduces TabFM: A Zero-Shot Foundation Model for Tabular Data

Editorial illustration: Google TabFM foundation model for zero-shot tabular data analysis

Google Research has published TabFM, a foundation model for tabular data that delivers zero-shot predictions in a single forward pass, without hyperparameter tuning or feature engineering. The model achieved top Elo scores on the TabArena benchmark and is available on Hugging Face and GitHub, with a planned integration into Google BigQuery.

🔴 🤖 Models June 29, 2026 · 3 min read

Meta: Brain2Qwerty v2 — Non-Invasive Thought-to-Text Decoding at 61% Accuracy, Without Surgical Implants

Editorial illustration: Brain2Qwerty v2 — non-invasive thought-to-text decoding at 61% accuracy, without surgical implants, without text or faces

Brain2Qwerty v2 is a Meta Research AI system that converts brain signals recorded outside the body — without surgery — into typed text at an average word-level accuracy of 61%, using MEG scanning. This is seven times higher than other non-invasive methods (8%). Training code and datasets have been released as open source.

🟡 🤖 Models June 29, 2026 · 2 min read

GitHub: Claude Opus 4.8 Fast Mode Arrives in Copilot Preview; Anthropic Retires Fast Mode for Opus 4.6

Editorial illustration: Claude Opus 4.8 fast mode arrives in Copilot preview; Anthropic retires fast mode for Opus 4.6, without text or faces

Claude Opus 4.8 fast mode is now in preview for GitHub Copilot users, delivering significantly faster output token generation while maintaining the model's intelligence level. Simultaneously, Anthropic is retiring fast mode for Opus 4.6 — consolidating fast mode capabilities on the sole remaining model.

🟢 🤖 Models June 29, 2026 · 2 min read

Allen Institute: DiScoFormer — One Transformer for Density and Score Across Distributions

Editorial illustration: DiScoFormer — one transformer for density and score across different distributions, without text or faces

DiScoFormer is an Allen Institute for AI (AI2) transformer model that estimates the density function (distribution density) and score function in a single forward pass — previously requiring separate models. It generalizes KDE to high dimensions and adapts to new distributions without retraining.

🟢 🤖 Models June 29, 2026 · 2 min read

arXiv:2606.28166: Tandem RL — Verifiable Rewards With More Readable Chain of Thought and Better Handoff to Smaller Models

Editorial illustration: 2606.28166: Tandem RL — verifiable rewards with a more readable chain of thought and better handoff, without text or faces

Tandem RL is a new language model training method that combines RLVR (reinforcement learning with verifiable rewards) with a tandem approach: a stronger model collaborates with a frozen weaker model during chain-of-thought generation. On Qwen3-4B it achieves comparable performance with significantly better readability and robustness when handing off to a smaller model.

🟡 🤖 Models June 28, 2026 · 2 min read

GitHub: MAI-Code-1-Flash, Microsoft's coding model, now generally available in Copilot Business and Enterprise plans

Editorial illustration: accelerated code flow through a development interface, without text or faces

MAI-Code-1-Flash is Microsoft's proprietary coding model that became generally available on June 26, 2026 for GitHub Copilot Business and Enterprise plans. Optimized for low latency and high-frequency agentic coding workflows, it is billed on a usage-based model according to the provider's pricing, and administrators must explicitly enable it in organization settings.

🟢 🤖 Models June 28, 2026 · 1 min read

arXiv:2606.26935: CoT training gains land in stronger action prediction, not deeper agent reasoning

Editorial illustration: a branching decision flow narrowing into a single clear path, without text or faces

A study by Jingyu Liu and colleagues (arXiv:2606.26935) shows that gains from chain-of-thought (CoT) training in LLM agents land in stronger direct action prediction rather than broader reasoning advantage. Later checkpoints revise the action less frequently, while masking supervision over action tokens improves out-of-domain generalization.

🟢 🤖 Models June 28, 2026 · 2 min read

arXiv:2606.26502: reasoning models spend more tokens on tasks they fail, opposite to humans who disengage

Editorial illustration: two effort curves diverging, one rising and one falling, without text or faces

A study by Han-yu Wang (arXiv:2606.26502) finds that large reasoning models (LRM) spend more tokens on tasks they ultimately get wrong than on those they solve correctly, the opposite of humans who disengage on harder tasks. The gap is large (Cohen's d 1.47–3.13 on the H-ARC benchmark), and all five tested models showed the reverse pattern from humans.

🔴 🤖 Models June 27, 2026 · 3 min read

OpenAI: GPT-5.6 Sol announced in preview — coding, science, and cybersecurity

Editorial illustration: abstract neural network visualization with code labels, molecular structure, and cybersecurity shield

GPT-5.6 Sol is OpenAI's announced next-generation model, currently in preview status (not generally available), with enhanced capabilities in coding, scientific reasoning, and cybersecurity, along with the most advanced safety stack to date.

🟡 🤖 Models June 27, 2026 · 2 min read

Anthropic: API rate limits raised — Sonnet and Haiku now match Opus across three tiers

Editorial illustration: chart with three API access tiers and upward arrows, abstract clouds and server racks

Anthropic has equalized API rate limits across all models — Sonnet and Haiku now share the same quotas as Opus at each of the three usage tiers (Start, Build, Scale). At the same time, fast mode for Claude Opus 4.7 is being deprecated, with removal scheduled for July 24.

🟡 🤖 Models June 27, 2026 · 2 min read

arXiv:2606.26836: Benchmarks miss 82% of AI models' real capabilities

Editorial illustration: bar chart showing benchmark gap between single-model evaluation and Capability Frontier measurement

Researchers showed that standard benchmarks — measuring only one model in one attempt — underestimate the true capabilities of LLMs by as much as 82%. By introducing the Capability Frontier framework, which uses Pareto optimality across 21 models and 16 benchmarks, the same accuracy is achievable at 85% lower cost.

🟡 🤖 Models June 27, 2026 · 2 min read

Google: Gemini Nano on Pixel is 50%+ faster with frozen multi-token prediction

Editorial illustration: smartphone chip diagram showing parallel token prediction paths on Pixel device

Google accelerated Gemini Nano inference on Pixel 9 and 10 by more than 50% using frozen multi-token prediction — a technique that generates an average of roughly 2 tokens per model pass, saving 130 MB of memory per instance with no change to output results.

🟢 🤖 Models June 27, 2026 · 2 min read

arXiv:2606.27288: When combining LLMs really helps — co-failure ceiling across 67 frontier models

Editorial illustration: diagram of accuracy ceiling for a group of AI models, abstract graphs without faces

A study of 67 frontier models from 21 providers introduces the concept of co-failure ceiling — the upper accuracy bound of an LLM ensemble determined by the rate at which all models fail on the same query. Results show that combining models rarely beats the single strongest model without query-level routing.

🟡 🤖 Models June 26, 2026 · 2 min read

Google Research: how thinking unlocks parametric knowledge in LLMs

Editorial illustration: stylized neural network pathways lighting up in sequence, abstract brain and data nodes, cool blue tones

Google Research reveals two mechanisms by which a reasoning trace improves retrieval of facts stored in model weights — computational buffer and factual priming — tested on Gemini 2.5 and Qwen3-32B.

🟢 🤖 Models June 26, 2026 · 2 min read

arXiv:2606.25325: OPPO — an RL framework that teaches AI to read emotions from voice, face, and text simultaneously

Editorial illustration: multimodal system with sound waves, a face frame, and text bubbles merging into a single emotion analysis

OPPO is a reinforcement learning system that trains omni-modal language models to simultaneously understand visual, acoustic, and textual emotion cues, suppressing cross-modal hallucinations and achieving SOTA results on two benchmark datasets.

🟡 🤖 Models June 25, 2026 · 2 min read

arXiv:2606.24510: RaDaR — specialized 32B reasoning LLM accelerates rare disease diagnosis in RCT

Editorial illustration: medical AI diagnostics, accuracy graphs, molecular structure and digital medical records

RaDaR is an open-source reasoning LLM with 32 billion parameters trained for rare disease diagnosis. In a randomized clinical trial it improved physician diagnostic accuracy by 21.44 percentage points versus internet search, with the ability to identify diagnoses in 61% of cases before clinical documentation.

🟡 🤖 Models June 25, 2026 · 2 min read

arXiv:2606.24014: RL training on health domain transfers alignment to 80%+ OOD benchmarks

Editorial illustration: neural network connections branching across multiple domains with alignment transfer arrows, abstract scientific visualization

Google Research researchers showed that RL training on beneficial properties such as truthfulness, fairness, and corrigibility improves performance on more than 80% of 50+ independent OOD benchmarks — including domains outside health on which the model was trained.

🟡 🤖 Models June 25, 2026 · 2 min read

Google: DiffusionGemma 26B — 4× faster text generation via diffusion approach

Editorial illustration: abstract depiction of parallel text streams forming from a diffusion cloud, digital style

DiffusionGemma is Google's 26B MoE model that generates text using a diffusion approach — in parallel rather than sequentially. It achieves more than 1,000 tokens per second on a single H100 GPU, up to 4× faster than standard autoregressive models, with a quality trade-off versus Gemma 4.

🟡 🤖 Models June 25, 2026 · 2 min read

Google: Gemini 3.5 Live Translate — speech-to-speech in 70+ languages in real time

Editorial illustration: audio waveforms connecting speech bubbles in multiple scripts across a globe, real-time translation concept

Google launched Gemini 3.5 Live Translate — a speech-to-speech translation system supporting 70+ languages and more than 2,000 language combinations in real time, with intonation preservation and SynthID watermark protection.

🟡 🤖 Models June 24, 2026 · 2 min read

arXiv:2606.23181: DART — training-free adaptive thinking in hybrid reasoning models

Editorial illustration: abstract branching diagram with two separate decision paths in a token network

DART is a routing method that decides without any training whether an AI model needs to think deeply or can respond immediately — reducing thinking token consumption by 15–69% while simultaneously improving accuracy by up to +22.5 points on code benchmarks.

🟡 🤖 Models June 24, 2026 · 2 min read

Mistral: OCR 4 — structured document extraction with bounding boxes in 170 languages

Editorial illustration: scanned paper document with labeled paragraphs and bounding boxes in various languages

Mistral OCR 4 is a new optical character recognition model that tops the OlmOCRBench leaderboard with 85.20 points, supports 170 languages, and delivers paragraph-level bounding boxes — all at a price of $4 per 1,000 pages.

🟡 🤖 Models June 24, 2026 · 2 min read

PyTorch/SGLang: DeepSeek-V4 Pro on NVIDIA GB300 — 5× higher throughput with the same interactivity

Editorial illustration: server rack with NVIDIA Blackwell GPU cards and a graph showing fivefold throughput growth

The PyTorch team and SGLang increased the serving throughput of the DeepSeek-V4 Pro model on NVIDIA GB300 architecture from around 2,200 to over 11,200 tokens per second per GPU between April and June 2026 — a fivefold improvement without sacrificing end-user interactivity.

🟡 🤖 Models June 21, 2026 · 2 min read

arXiv:2606.20560: DiffusionGemma as interpretable as Gemma 4 — 28.6× gap reduced to 1.1×

Editorial illustration: DiffusionGemma as interpretable as Gemma 4 — 28.6× gap reduced to 1.1×

DiffusionGemma is Google's diffusion language model operating in continuous latent space. A study by 13 authors led by Neel Nanda shows that the initial opacity is 28.6× greater than Gemma 4, but an interpretable token bottleneck narrows that gap to just 1.1×.

🟢 🤖 Models June 21, 2026 · 2 min read

arXiv:2606.20543: Spatially Speculative Decoding accelerates image generation 13.3×

Editorial illustration: Spatially Speculative Decoding accelerates image generation 13.3×

SSD (Spatially Speculative Decoding) is a new method that simultaneously predicts the horizontal and vertical neighbor of a pixel in autoregressive image generation, achieving up to 13.3× speedup without any loss of visual quality on the DPG-Bench and GenEval benchmarks.

🟢 🤖 Models June 21, 2026 · 2 min read

arXiv:2606.20561: TimeProVe Reduces Long-Video Reasoning Inference Costs by 93%

Editorial illustration: TimeProVe reduces long-video reasoning inference costs by 93%

TimeProVe is a framework that accelerates VLM inference over long videos by introducing a two-stage propose-then-verify approach. It reduces calls to expensive models by 75% and total inference cost by 93%, while outperforming the strongest competitor by 7.3 percentage points on the new OpenTSUBench benchmark.

🟢 🤖 Models June 21, 2026 · 2 min read

arXiv:2606.20008: VIMPO — Critic-Free Reinforcement Learning Beats GRPO on MATH-500 and AIME

Editorial illustration: VIMPO — critic-free reinforcement learning beats GRPO on MATH-500 and AIME

VIMPO is a new reinforcement learning method for LLM reasoning that derives an implicit value function from KL-regularized RL — without a separate critic network. It outperforms GRPO on four mathematical benchmarks including AIME 2024 and AIME 2025, with gains that remain stable even under noisy reward conditions.

🟡 🤖 Models June 20, 2026 · 2 min read

arXiv:2606.20333: SoftSkill Compresses Skill Documents into 32 Latent Tokens and Boosts LiveMath by 42.1 Points

Editorial illustration: a long document compressed into a few glowing latent tokens

SoftSkill is a method described in the paper arXiv:2606.20333 that converts skill documents, such as Markdown SKILL.md files, into compact continuous latent objects that guide model behavior without modifying the base model. Instead of hundreds or thousands of instruction tokens, SoftSkill uses 32 virtual tokens and achieves an improvement of 42.1 percentage points on the LiveMath task compared to operating without the skill.

View full archive →