Enkrypt AI Risk Scores Are Now Live on Kilo Code


Kilo now shows an Enkrypt AI risk score on every scored model card at kilo.ai/leaderboard, right next to completion rate and cost per attempt, the same evaluation behind Enkrypt's public AI Safety Leaderboard.
β
Why This Matters
β
Choosing a model for a coding task means weighing speed and cost against whether its output is safe to deploy. Kilo has always surfaced the first part. How a model handles jailbreak attempts, bias, toxicity, and insecure code is what Enkrypt AI's evaluation measures, not just whether a model carries risk, but where: which attack categories succeed, how often, and under what conditions.
β
That data has always lived on Enkrypt AI's own leaderboard. What hasn't existed until now is visibility into it at the point a model actually gets selected. Kilo Code is where a lot of those selections happen, and a risk score on the same card as price and completion rate means Enkrypt's evaluation shapes the decision, rather than reviewing it after the fact.
β

β
What's Being Measured
β
This is the same evaluation Enkrypt AI has run across 200+ models on its own leaderboard, independent of anything Kilo tracks internally: jailbreak resistance, bias, toxicity, and insecure code generation, among other categories, combined into a single 0 to 100 score, lower is safer. No separate step for Kilo Code users, it's already on the card.
β
Insecure code generation is the one worth paying attention to for anyone running agents with less oversight than a standard code review. A model can complete tasks quickly and read fine in a diff while still being more likely to introduce a vulnerability a human reviewer would have caught.
β

β
What the Data Shows
β
As of this writing, Claude Sonnet 5 holds the lowest risk score on the board (9.2) at $36.19 per attempt. GPT-5.5 completes more tasks (74.2% vs. 59.6%) but costs about twice as much per attempt and carries roughly double the risk score. Cost, completion rate, and risk score don't move together, this is the first place all three sit side by side for the same model.
β

β
What's Next
β
Open the All Models tab and check where your model lands before you commit to one, especially for anything running unsupervised. More on how this evolves is coming soon. Stay tuned!
Frequently Asked Questions
No. Kilo's ranking is based on real usage and completion rate. Risk score is a separate data point.
No. If you use Kilo Code, it's already on the model card.
Jailbreak resistance, bias, toxicity, and insecure code generation are among the categories Enkrypt AI tests, rolled into one 0 to 100 number. A deeper breakdown is coming in a follow-up post.
Enkrypt AI evaluates models individually; coverage expands over time. Check the All Models tab for current status.
Yes, Anaconda acquired Kilo Code in July 2026 and Enkrypt AI in August. This is the first place both show up together.
.avif)



