FAR.AI has introduced a public leaderboard comparing AI model safeguards, suggesting that resilience against misuse varies significantly even among today’s leading frontier systems.
FAR.AI has launched an AI Security Leaderboard that evaluates how effectively leading AI models withstand attempts to bypass their built-in safety measures. The nonprofit’s findings suggest that while some frontier models resisted every tested attack, others proved considerably more vulnerable under the same evaluation conditions. As artificial intelligence becomes more capable and widely deployed, independent assessments of model safety are emerging as an increasingly important part of the broader discussion around responsible AI development.
Public attention often focuses on the capabilities of advanced AI systems, but their safeguards have become equally significant as the technology expands into sensitive domains. Organizations, governments, and researchers are increasingly concerned with whether models can reliably refuse requests related to cyberattacks, hazardous materials, or other forms of misuse. This has led to growing interest in standardized testing methods that measure not only what AI systems can do, but also how consistently they prevent harmful behavior.
According to FAR.AI, the evaluation tested four frontier models across chemical, biological, radiological, nuclear, explosive, and cybersecurity scenarios using more than 1,500 automated and expert-guided attack attempts. The organization reported that it identified hundreds of what it describes as “universal jailbreaks” affecting Grok 4.5 and Gemini 3.1 Pro, while it found none for Claude Fable 5 or GPT-5.6 Sol during the same testing process. FAR.AI also estimated a substantial difference in the effort required to discover successful attacks, reporting that effective jailbreaks were found for some models at relatively low cost, whereas comparable attempts against others remained unsuccessful throughout the evaluation. Alongside the leaderboard, the organization published a proposed minimum standard for AI safeguards intended as a baseline rather than a certification of security.
The report underscores a broader shift in AI governance, where safety claims are increasingly expected to be supported by transparent, repeatable testing rather than company assertions alone. Independent benchmarks can provide policymakers, enterprise buyers, and researchers with a clearer basis for comparing systems, particularly as AI models are integrated into critical business and public sector applications. At the same time, FAR.AI emphasizes that resisting the attacks included in this evaluation should not be interpreted as proof that any model is completely secure.
FAR.AI’s initiative reflects the growing maturity of AI safety as a field of independent evaluation. As frontier models continue to evolve, ongoing public testing and comparative benchmarks may become an important complement to performance rankings, helping shift attention from capability alone toward the resilience and reliability of AI systems deployed in the real world.