Guide /
Fine-tuning a small language model with LoRA
A practical starting point for supervised fine-tuning: define the task, prepare data, train an adapter, and evaluate the result.
Fine-tuning is most useful when the behavior you want is specific enough to describe and measure. Before changing model weights, write down the task, the expected input and output, and a small set of examples that the base model currently gets wrong. A good baseline may be a carefully written prompt. Training should improve on that baseline on data the model has not seen.
Start with a narrow task
Suppose the task is to turn short technical questions into concise, evidence-based answers. Collect examples with a consistent format. Remove duplicates, check that the answers are correct, and keep the validation and test examples separate before training. If several examples come from the same document or person, split by that source rather than by individual row. Otherwise, near-duplicates can make the evaluation look better than it is.
Use a model whose license permits your intended use. Read its model card for the expected chat template and tokenizer. A small public model makes the first experiment cheaper and easier to debug; the exact model is less important than a clean evaluation setup.
Train an adapter
LoRA keeps the base weights fixed and trains small adapter matrices. In the Hugging Face stack, SFTTrainer can take a LoraConfig through peft_config. The following is the shape of an experiment, with the dataset deliberately left as your own reviewed training set:
from peft import LoraConfig
from trl import SFTConfig, SFTTrainer
adapter = LoraConfig(
r=16,
lora_alpha=32,
lora_dropout=0.05,
bias="none",
task_type="CAUSAL_LM",
target_modules=["q_proj", "v_proj"],
)
config = SFTConfig(
output_dir="./outputs/first-adapter",
num_train_epochs=1,
per_device_train_batch_size=2,
learning_rate=2e-4,
)
trainer = SFTTrainer(
model="Qwen/Qwen2-0.5B",
args=config,
train_dataset=train_dataset, # Your reviewed, correctly formatted dataset
peft_config=adapter,
)
trainer.train()
This is a template, not a claim of a completed experiment. Dataset formatting, chat templates, package versions, and hardware requirements must be checked for the model you actually choose. Start with a few examples to verify tokenization and loss before committing to a full run.
Evaluate the behavior, not only the loss
Compare the base model and the adapter on the same held-out examples. Look at task-specific correctness, unsupported claims, failure cases, and whether the model follows the required format. Keep a record of prompts, model revision, data version, adapter settings, and random seed. If the task is safety-sensitive, add expert review and appropriate safeguards before considering use beyond an experiment.
Training loss tells you that optimization happened. It does not tell you whether the system became useful, reliable, or safe. The evaluation design makes that distinction visible.
References: Hugging Face TRL: PEFT integration, Hugging Face PEFT: LoRA.