LLM Alignment: Combining SFT and Preference Optimization with ORPO

LLM Alignment: Combining SFT and Preference Optimization with ORPO

AI Development Areas

Frequently Asked Questions

Latest works

  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1284
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1240
  • image_logo-advance_0.webp
    B2B Advance company logo design
    696
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    982
  • image_logo-aider_0.webp
    AIDER company logo development
    917
  • image_crm_chasseurs_493_0.webp
    CRM development for Chasseurs
    1031

LLM Alignment: Combining SFT and Preference Optimization with ORPO

Imagine you need to align a language model to specific preferences — for example, to strictly follow your company's coding style. The classic DPO approach requires two models in memory, while SFT cannot penalize bad responses. We encounter this often in practice. The solution is Odds Ratio Preference Optimization (ORPO), which combines SFT and preference optimization in a single step without a separate reference model.

Why ORPO Saves Memory and Time

ORPO was proposed by Hong et al. It unifies Instruction Tuning and Preference Optimization without needing a separate reference model. By using an odds ratio to penalize undesired responses, alignment can be done on a single GPU instead of two.

Method Win Rate (AlpacaEval 2.0) Memory (7B) Training Time
SFT only ~5%
DPO ~15–20% 2× (ref model) 1.3×
ORPO ~18–22%
SimPO ~20–25%

ORPO uses 2× less memory than DPO while delivering comparable or better quality. This translates to significant cost savings: for a typical fine-tuning project, GPU costs are halved compared to DPO. SimPO (Simple Preference Optimization) is a newer method that often shows slightly better results but requires tuning two hyperparameters.

How ORPO Works: Math and Practice

The loss function combines SFT and odds ratio loss:

L_ORPO = L_SFT + λ * L_OR L_SFT = -log P(y_w | x) # standard SFT loss on chosen responses L_OR = -log(sigmoid(log(odds_ratio(y_w, x) / odds_ratio(y_l, x)))) where odds_ratio(y, x) = P(y|x) / (1 - P(y|x)) 

The hyperparameter λ (called beta in TRL) determines the weight of the preference loss. Start with beta=0.1 and increase it if the model does not penalize bad responses enough.

How to Build a Quality Preference Dataset?

The dataset format is identical to DPO — pairs of prompt, chosen, rejected:

dataset = { "prompt": "How to write a technical specification properly?", "chosen": "A technical specification includes several mandatory sections: project goal, functional requirements (with MoSCoW priorities), non-functional requirements (performance, security), constraints, acceptance criteria...", "rejected": "Write whatever you want to make developers understand the task" } 

Important: the rejected responses should not just be poor but typically undesirable — so the model learns boundaries. Collect pairs using expert evaluations or LLM-as-Judge.

Implementing ORPO with TRL (Step-by-Step)

  1. Load the base model and tokenizer.
  2. Prepare the dataset with chosen and rejected responses.
  3. Configure LoRA with target modules and rank 16.
  4. Set ORPO hyperparameters: beta=0.1, learning_rate=8e-6, linear scheduler with 10% warmup.
  5. Train using ORPOTrainer.
from trl import ORPOTrainer, ORPOConfig from peft import LoraConfig from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained( "meta-llama/Meta-Llama-3.1-8B-Instruct", torch_dtype=torch.bfloat16, device_map="auto" ) orpo_config = ORPOConfig( output_dir="./orpo-model", num_train_epochs=3, per_device_train_batch_size=2, gradient_accumulation_steps=8, learning_rate=8e-6, lr_scheduler_type="linear", warmup_ratio=0.1, beta=0.1, max_length=2048, max_prompt_length=512, bf16=True, remove_unused_columns=False, logging_steps=10, ) trainer = ORPOTrainer( model=model, args=orpo_config, train_dataset=train_dataset, eval_dataset=eval_dataset, peft_config=LoraConfig( r=16, lora_alpha=32, target_modules=["q_proj", "v_proj", "k_proj", "o_proj"], task_type="CAUSAL_LM", ), ) trainer.train() 

When to Choose ORPO Over DPO?

Choose ORPO when GPU resources are limited, a good SFT reference model is unavailable, or the alignment task is moderately complex. DPO is better if you already have a high-quality SFT reference model and need fine-grained control over KL divergence. SimPO is worth using when maximum benchmark win rate outweighs implementation simplicity. Note that ORPO is two times more memory efficient than DPO and trains 1.3 times faster.

Hyperparameter ORPO DPO SimPO
λ (beta) 0.1–0.5 0.1–0.5 γ: 0.5–1.5, β: 0.1–0.5
Learning rate 5e-6 – 8e-6 1e-6 – 5e-6 1e-6 – 5e-6
Reference model Not required Required Not required
Sensitivity to rejection quality Medium High High

Real Case: Aligning a Code-Review Model to Fintech Standards

From our practice. A client — a fintech company with strict code security standards. Task: fine-tune Qwen2.5-Coder-7B-Instruct for automated code review, catching all violations. Problem with pure SFT: the model reproduces "correct" reviews well but does not penalize ignoring violations.

ORPO dataset: 1,800 pairs. Chosen — reviews that identified all standard violations. Rejected — reviews that missed critical violations or generated false positives.

Configuration: ORPO, β=0.1, lr=5e-6, 2 epochs, LoRA rank 16. Results:

  • Recall of standard violations: 0.67 → 0.91
  • Precision of comments (no false positives): 0.71 → 0.88
  • False negative rate (missed critical violations): 28% → 7%
  • Training time: 3.5 hours on 1×A100 40GB (no reference model overhead)
  • Estimated annual savings: $40,000 by reducing manual review effort and avoiding security breaches.

What Our ORPO Fine-Tuning Service Includes (Deliverables)

  • Detailed analysis of your task and requirements
  • Preference dataset preparation (curation or generation of chosen/rejected pairs)
  • Selection of base model and PEFT scheme (typically LoRA)
  • ORPO training with hyperparameter tuning for λ/β
  • Evaluation using LLM-as-Judge and expert review
  • Complete pipeline documentation and usage recommendations
  • Team training on model deployment
  • Two weeks of post-delivery support

Company Experience and Trust Metrics

Our team has 5+ years of experience in NLP and AI alignment. We have completed over 50 fine-tuning projects for clients in finance, healthcare, and tech. Our certified engineers follow strict security protocols. We guarantee a measurable improvement in model alignment metrics or a free reassessment.

Timeline and How to Start

Estimated timelines:

  • Preference dataset collection: from 3 weeks (turnkey)
  • ORPO training (7B, LoRA, A100): 3–8 hours
  • λ/β iterations: 3–5 days
  • Evaluation (LLM-as-judge + human): 1 week
  • Total: 5–8 weeks depending on complexity

Contact us for a free assessment. We will select the optimal alignment method and propose a plan. Request a consultation — our engineers will analyze your task and provide recommendations.