
Lyzr Open FineTuner
How the data preparation and fine-tuning pipeline works
Documents become a clean training set through two de-dup gates, train through LoRA and then DPO, and reach agents only after the evals pass. Two live triggers keep it running. A change to the source documents restarts data preparation, and a failed scheduled eval starts refinetuning.
De-dup gate, two of themLive triggerRefinetuning loop when evals failOutput of training
Prepare
- Documents trigger
- Source corpus. Any change here triggers the data preparation pipeline. Duplicate files are dropped before chunking.
- Chunking
- Documents split into passages.
- Pair generation
- Each chunk yields two kinds of training pairs.
Build dataset
- QA pairs
- A question and its reference answer. Used for supervised training.
- Preference pairs
- A prompt with a chosen and a rejected answer. Used for DPO.
- Training dataset
- Both sets merged after the second de-dup, at pair level.
Train
- LoRA
- Supervised fine-tune with low-rank adapters on the QA pairs.
- DPO
- Direct preference optimization on the preference pairs, starting from the LoRA adapter.
Evaluate
- Fine-tuned model
- The output of LoRA followed by DPO.
- Evals trigger
- Continuous evals run on a schedule against the continuous test cases.
- Continuous test cases
- The benchmark every model has to clear before it serves.
- Refinetuning loop
- Kicks in when evals fail. Fine-tunes a new or existing model on the latest data, re-runs the evals, and ships only models that clear the benchmarks.
Serve
- Inference
- Only a model that passed its evals is served.
- Agents
- Agents call the served model.