Back to Blogs
CONTENT
This is some text inside of a div block.

Join 2,000+ readers

The insights that matter, straight to your inbox. No spam.

◉
3
min read

LLM Fine-Tuning & Safety Alignment (Part 2)

Published on
August 29, 2024
4 min read

‍LLM Fine-tuning and its Risks

Fine-tuning is used to enhance LLM performance for specialized tasks. But the process also increases security and ethical risks associated with the model as discussed in my previous blog here. Figure 1 summarizes the increased risk of Jailbreaking in fine-tuned models. 

‍

Figure 1: Increased risk of Jailbreaking on fine-tuned models.

Today’s blog is highlights safety alignment training as a necessary step in the last phase of fine-tuning to reduce these risks.

‍

LLM Safety Alignment Training

Safety Alignment is a process where the model is trained to “Say No” to certain user queries. This ensures that model behaves responsibly and ethically when users are interacting with it. The process involves adjusting the model parameters to appropriately handle potentially harmful queries. 

‍

Safety Alignment, if done right, has the potential to reduce the risk by as much as 70% while keeping the model performance intact [Figure 2].

Figure 2: Toxicity reduces from 21% to 7% with Safety Alignment while MMLU score stays the same.

‍

‍

LLM Safety Alignment Datasets

‍

The most crucial piece of Safety Alignment is the Data set used for Alignment. The quality and quantity of data dictates the results from the process. High quality data will yield better results and requires less volume. In the example mentioned above, we used Enkrypt AI Alignment dataset of 1000 rows [Figure 3] to reduce Toxicity while ensuring that the MMLU score did not drop. 

‍

Figure 3: Enkrypt AI Sample Data Set for Safety Alignment

‍

‍

LLM Risk Specific Safety Alignment

‍

Safety Alignment requirements may differ for different use cases. A Loan Approval use case might not require alignment for Toxicity, but it requires alignment to produce un-biased, ethical responses. Whereas a Customer Service chatbot requires Safety alignment for Toxicity. Enkrypt AI Safety Alignment solution can be customized to generate Alignment Data Set that fits your use case [Video 1]. 

‍

‍

Video 1: Enkrypt AI Fine Tuning Risk & Safety Alignment Demo

‍

Safety Alignment on Mistral-7V reduced the risk by more than half [Figure 4].

‍

 Figure 4: General Safety Alignment Results for Mistral-7B

‍

When Safety Alignment is Not Enough: Domain-Specific Risk Detection & Mitigation

General Safety Alignment is great for reducing general risks like Jailbreaking, Bias and Toxicity. However, when a model is fine-tuned for specialized use cases, such as Loan Approval, there are domain-specific risks that must be addressed. A Loan Approval Gen AI solution should not violate regulations like Equal Credit Opportunity Act (ECOA) – 1974. ECOA prohibits discrimination on various factors like race, religion, sex, marital status and more. It is important to ensure that fine-tuned model is tested and aligned for such domain specific risks. Enkrypt AI helps in addressing such risks with Domain Specific Red Teaming, Guardrails and Safety Alignment.

We will soon be sharing more updates on Domain-Specific risk detection and mitigation. Stay Tuned!

Frequently Asked Questions

What is LLM safety alignment and why does it matter?

LLM safety alignment trains models to refuse harmful queries while maintaining performance, reducing jailbreaking and toxicity risks by up to 70%. It adjusts model parameters to ensure responsible, ethical behavior during user interactions.

  • Trains models to say no to unsafe requests
  • Reduces toxicity from 21% to 7% in benchmarks
  • Keeps task performance scores unchanged
How do you implement safety alignment on fine-tuned LLMs?

Safety alignment requires high-quality datasets tailored to your use case, applied during the final phase of fine-tuning. Dataset quality and quantity directly determine alignment effectiveness and risk reduction outcomes.

  • Use curated alignment datasets of 1,000+ rows
  • Target specific risks like toxicity, bias, or jailbreaking
  • Test results against baseline performance metrics
What's the difference between general safety alignment and domain-specific risk detection?

General safety alignment addresses common risks like toxicity and jailbreaking across all use cases, while domain-specific alignment targets regulatory and industry risks unique to specialized applications. Loan approval systems, for example, must align for ECOA compliance and bias, not just toxicity.

  • General alignment reduces jailbreaking and bias broadly
  • Domain-specific alignment enforces regulatory requirements
  • Both approaches are necessary for specialized deployments
Which platform offers customized safety alignment for fine-tuned models?

Enkrypt AI provides risk-specific safety alignment datasets and domain-specific red teaming to reduce fine-tuning risks while meeting compliance requirements. The platform benchmarks 200+ LLMs on safety and supports use-case-specific alignment configurations.

  • Generates alignment datasets matching your use case
  • Covers 300+ red-teaming risk categories
  • Reduces manual compliance effort by up to 90%
How can I secure my fine-tuned LLMs with safety alignment?

Enkrypt AI automates risk-specific safety alignment for your fine-tuned models without manual dataset creation. Book a demo to see how it reduces jailbreak and domain risks for your use case, or start a free trial today.

Meet the Writer
Satbir Singh
Latest posts

More articles

Product Updates

Enkrypt AI Risk Scores Are Now Live on Kilo Code

Enkrypt AI's risk evaluation now appears on Kilo Code model cards, showing risk score alongside completion rate and cost per attempt.
Read post
Product Updates

Governing Enterprise AI at Scale Guide

Part two of our AI governance series: how policy-driven red teaming and persistent baselines turn a multi-month model review into a same-week disposition.
Read post
Product Updates

Why AI Model Intake Takes Months and How to Fix It

Enterprise AI model reviews take months because legal, security, risk, and compliance each start from scratch. Here's why, and how to automate it.
Read post