🟡 🛡️ Security Published: · 2 min read ·

Sakana AI: Fugu-Cyber scores 86.9% on CyberGym and 72.1% on CTI-REALM, comparable to GPT-5.5-Cyber

Abstract depiction of a multi-agent AI system analyzing cybersecurity vulnerabilities

Sakana AI has introduced Fugu-Cyber, a multi-agent orchestration system specialized in cybersecurity that behaves as a single model. It scores 86.9% on the CyberGym vulnerability-analysis benchmark and 72.1% on the CTI-REALM threat-intelligence translation benchmark, comparable to GPT-5.5-Cyber and Mythos-Preview. It is available via API with manual access approval.

🤖

This article was generated using artificial intelligence from primary sources.

What is Fugu-Cyber and why does it act as a single model?

Sakana AI has introduced Fugu-Cyber, a multi-agent orchestration system — an architecture in which multiple specialized AI agents collaborate to solve a task — specialized exclusively for cybersecurity. Although it internally coordinates multiple agents, the system behaves and responds to the user as a single, unified model, with no visible transitions between individual components.

Results on the CyberGym and CTI-REALM benchmarks

Fugu-Cyber achieves 86.9% success on the CyberGym benchmark, which measures the ability to analyze vulnerabilities in complex, real-world code. On the CTI-REALM benchmark, which tests translating threat intelligence — intelligence data on threats gathered from various sources — into concrete detection rules for security systems, the model scores 72.1%. Sakana AI states that both results are comparable to those achieved by GPT-5.5-Cyber and Mythos-Preview, two leading competing models specialized in a similar segment of cybersecurity.

Who is allowed to use Fugu-Cyber?

Access to the model goes through the sakana.ai/fugu API endpoint, but is limited exclusively to subscribers of Sakana’s Token Plan. Every access request undergoes manual approval, and the user must provide a verified contact before receiving an API key. This approval model follows the company’s Acceptable Usage Policy, which explicitly prohibits offensive misuse of the tool — for example, using the system to develop attacks rather than defenses.

Why does comparability with GPT-5.5-Cyber matter?

Results on both benchmarks show that Fugu-Cyber does not lag significantly behind competing specialized models despite its different, multi-agent architecture. For security teams, this means an additional option in the market of tools for vulnerability analysis and threat-intelligence processing, with stricter access controls than has typically been the case with API offerings so far. The comparison with GPT-5.5-Cyber and Mythos-Preview shows that a multi-agent architecture, despite the added complexity of orchestration, does not have to mean lower performance on specialized cybersecurity tasks — on the contrary, coordinating multiple agents can be an advantage on tasks that require both code analysis and interpreting intelligence data in the same step.

Frequently Asked Questions

What is Fugu-Cyber and how does it work?
Fugu-Cyber is a multi-agent orchestration system from Sakana AI specialized for cybersecurity tasks, which internally coordinates multiple specialized agents but presents itself to the user as a single, unified model.
How successful is Fugu-Cyber on security benchmarks?
Fugu-Cyber achieves 86.9% success on the CyberGym benchmark for analyzing vulnerabilities in complex code and 72.1% on the CTI-REALM benchmark for translating threat intelligence into detection rules, comparable to the results of the GPT-5.5-Cyber and Mythos-Preview models.
Who can access the Fugu-Cyber model?
Access is limited to subscribers of Sakana's Token Plan, with manual approval of every request and a verified contact required, in accordance with an Acceptable Usage Policy that prohibits offensive misuse of the tool.

📬 AI news in your inbox

A daily digest built your way — pick topics, sources and cadence. One-click unsubscribe.