Qwen 3.5–0.8B
84.6 vs. 83.0 for GPT 5.6-sol.
Shopify’s internal quality score.
Specialists trained on your traffic.
We do the work.
did not outperform
GPT-4 on their task.
LoRA Land: 310 fine-tuned models across 31 tasks, using 2B/7B-class models.
84.6 vs. 83.0 for GPT 5.6-sol.
Shopify’s internal quality score.
96% vs. 90% accuracy for o3.
64× lower inference cost.
53.7 vs. 38.3 for Llama-2 70B.
HumanEval + MBPP average.
82.5 vs. 64.7 for Mixtral 8×7B.
GSM8K · 8-shot reasoning.
55.5 vs. 9.3 for GPT-4o-0513.
AIME 2024 · pass@1.
56.0 vs. 41.0 for GPT-4o.
BFCL v3 multi-turn · FC mode.
61.06 vs. 59.74 for Llama 3.2 3B.
Average of seven benchmarks.
Beat Gopher 280B on 9 of 16
Pile language-modeling datasets.
make heavy use of fine-tuning.
Most still rely on base models.
Frontier-lab researcher
total compensation.
A better base model can arrive
before your investment
pays back.
Re-evaluate. Retrain. Redeploy.
The work keeps coming back.
There's a very high fixed cost to fine-tuning that you have to amortize over a large volume.
For Flo, the economics depend on volume and a narrow, critical task.

Lower cost and p95 latency. Continuously maintained by our ML team.
Once per agent/task.
They define correct labeling.
Training, evaluation
and deployment included.
Match or beat it
on your agreed task.
Our team uses it to train, evaluate, serve and improve your models.
Start from real agent inputs and outputs.
[PERSON][EMAIL][PHONE]Remove sensitive fields before training.
Choose the right base model and size.
Match compute to training and serving.
Train specialists on your agent’s work.
Align evaluations with human judgment.
Serve specialists in your agent workflow.
Quality · Cost · Latency
Improve as models and traffic change.
Mistral 7B · task-specific evaluation
Qwen 14B specialist vs. OpenAI o3
Median training-set size across all 31 LoRA Land tasks.
14,041 examples per training set (median) · 72% of fine-tuned models beat GPT-4 in LoRA Land.
$0.001 per SLM specialist call

Second-time founder.
Hands-on experience building custom models for real workflows.

Founder, Vidia
Stanford MSx

Founder, Zipper
Acquired by CRM&BONUS

Machine Learning Engineer
Nubank
With specialists · per year
A new generation is still entering.