Back to Blogs
CONTENT
This is some text inside of a div block.

Join 2,000+ readers

The insights that matter, straight to your inbox. No spam.

β—‰
3
min read

Enkrypt AI Risk Scores Are Now Live on Kilo Code

Published on
September 10, 2026
4 min read

Kilo now shows an Enkrypt AI risk score on every scored model card at kilo.ai/leaderboard, right next to completion rate and cost per attempt, the same evaluation behind Enkrypt's public AI Safety Leaderboard.

‍

Why This Matters

‍

Choosing a model for a coding task means weighing speed and cost against whether its output is safe to deploy. Kilo has always surfaced the first part. How a model handles jailbreak attempts, bias, toxicity, and insecure code is what Enkrypt AI's evaluation measures, not just whether a model carries risk, but where: which attack categories succeed, how often, and under what conditions.

‍

That data has always lived on Enkrypt AI's own leaderboard. What hasn't existed until now is visibility into it at the point a model actually gets selected. Kilo Code is where a lot of those selections happen, and a risk score on the same card as price and completion rate means Enkrypt's evaluation shapes the decision, rather than reviewing it after the fact.

‍

‍

What's Being Measured

‍

This is the same evaluation Enkrypt AI has run across 200+ models on its own leaderboard, independent of anything Kilo tracks internally: jailbreak resistance, bias, toxicity, and insecure code generation, among other categories, combined into a single 0 to 100 score, lower is safer. No separate step for Kilo Code users, it's already on the card.

‍

Insecure code generation is the one worth paying attention to for anyone running agents with less oversight than a standard code review. A model can complete tasks quickly and read fine in a diff while still being more likely to introduce a vulnerability a human reviewer would have caught.

‍

‍

What the Data Shows

‍

As of this writing, Claude Sonnet 5 holds the lowest risk score on the board (9.2) at $36.19 per attempt. GPT-5.5 completes more tasks (74.2% vs. 59.6%) but costs about twice as much per attempt and carries roughly double the risk score. Cost, completion rate, and risk score don't move together, this is the first place all three sit side by side for the same model.

‍

‍

What's Next

‍

Open the All Models tab and check where your model lands before you commit to one, especially for anything running unsupervised. More on how this evolves is coming soon. Stay tuned!

Frequently Asked Questions

Does the risk score affect Kilo's ranking?

No. Kilo's ranking is based on real usage and completion rate. Risk score is a separate data point.

Do I need an Enkrypt AI account to see it on Kilo?

No. If you use Kilo Code, it's already on the model card.

What exactly gets scored?

Jailbreak resistance, bias, toxicity, and insecure code generation are among the categories Enkrypt AI tests, rolled into one 0 to 100 number. A deeper breakdown is coming in a follow-up post.

Why isn't every model scored yet?

Enkrypt AI evaluates models individually; coverage expands over time. Check the All Models tab for current status.

Is this related to the acquisitions?

Yes, Anaconda acquired Kilo Code in July 2026 and Enkrypt AI in August. This is the first place both show up together.

Meet the Writer
Sheetal J
Latest posts

More articles

Product Updates

Governing Enterprise AI at Scale Guide

Part two of our AI governance series: how policy-driven red teaming and persistent baselines turn a multi-month model review into a same-week disposition.
Read post
Product Updates

Why AI Model Intake Takes Months and How to Fix It

Enterprise AI model reviews take months because legal, security, risk, and compliance each start from scratch. Here's why, and how to automate it.
Read post
Company News

What It Means To Secure AI on Your Own Terms: Anaconda Acquires Enkrypt AI

Anaconda acquires Enkrypt AI, embedding AI security, governance, and compliance controls across its platform to help enterprises secure agents at scale.
Read post