Nest

Specialists trained on your traffic.
We do the work.

Nest01
THE SHIFT TO PRODUCTION

Agents are already in production.

SHARE OF RESPONDENTS WITH PRODUCTION AGENTS
SurveyIllustrative: ~12.4% growth / year
~51%
2024
57.3%
2025
64.4%
2026
72.3%
2027
81.3%
2028
Scenario repeats the relative growth from 2024 → 2025.
LangChain 2024 · 2025 survey · Different samples. No comparable 2023 datum found. 2026–2028 are a scenario, not observed data.
Nest02
MID-SIZED COMPANIES · 100–2,000 EMPLOYEES

Three problems stand between
an agent and scale.

34.4%Quality
18.4%Latency
16%Cost
010203040%
LangChain · State of Agent Engineering, Dec 2025 · Company-size breakdown. 34.4% + 18.4% + 16% = 68.8%, rounded.
Nest03
SPECIALIZATION WORKS
72%

of fine-tuned models beat GPT-4 on-task,
with up to 30× potential compute savings.*

28%

did not outperform
GPT-4 on their task.

LoRA Land: 310 fine-tuned models across 31 tasks, using 2B/7B-class models.

LoRA Land, 2024 · 224/310 beat GPT-4-0613. *NVIDIA, 2025: separate efficiency estimate for 7B vs. 70–175B; not a measured API-price reduction.
Nest04
Proven in practice

Small models. Bigger results.

Published comparisons, from research to production.
~0.8B
Total parameters

Qwen 3.5–0.8B

Buyer profiles · Sep 2026

84.6 vs. 83.0 for GPT 5.6-sol.
Shopify’s internal quality score.

~14.7B
Total parameters

ART·E · Qwen 2.5–14B

Email research · Apr 2025

96% vs. 90% accuracy for o3.
64× lower inference cost.

~2.7B
Total parameters

Phi-2

Code generation · Dec 2023

53.7 vs. 38.3 for Llama-2 70B.
HumanEval + MBPP average.

~3.8B
Total parameters

Phi-3-mini

Math reasoning · 2024

82.5 vs. 64.7 for Mixtral 8×7B.
GSM8K · 8-shot reasoning.

~7B
Total parameters

R1-Distill-Qwen-7B

Competition math · Jan 2025

55.5 vs. 9.3 for GPT-4o-0513.
AIME 2024 · pass@1.

~3B
Total parameters

xLAM-2–3B

Multi-turn tool use · Apr 2025

56.0 vs. 41.0 for GPT-4o.
BFCL v3 multi-turn · FC mode.

~1.5B
Total parameters

Hymba-1.5B-Base

General benchmarks · Nov 2024

61.06 vs. 59.74 for Llama 3.2 3B.
Average of seven benchmarks.

~7.5B
Total parameters

RETRO

Retrieval-augmented · Dec 2021

Beat Gopher 280B on 9 of 16
Pile language-modeling datasets.

B = billion · Approximate reported model sizes · Benchmark scopes, model versions and dates are shown for each case.
Nest05
THE ADOPTION GAP
13.8%30.5%55.7%
Heavy useExperimentedNever fine-tuned

Few teams make
fine-tuning a
production habit.

13.8%

make heavy use of fine-tuning.
Most still rely on base models.

LangChain · State of Agent Engineering, Dec 2025 · “Experimented” respondents mainly use base models.
Nest06
WHY TEAMS HESITATE

The cost of owning the whole process.

Expensive to build

$600K+ / year

Frontier-lab researcher
total compensation.

The frontier moves

A better base model can arrive
before your investment
pays back.

Expensive to maintain

Re-evaluate. Retrain. Redeploy.
The work keeps coming back.

KORE1, 2026 · Frontier-lab researcher compensation includes equity; not a market-wide salary or minimum project staffing cost.
Nest07
Why fine-tuning needs scale
There's a very high fixed cost to fine-tuning that you have to amortize over a large volume.

For Flo, the economics depend on volume and a narrow, critical task.

Lindy
Flo CrivelloFounder & CEO, Lindy · June 2025
Flo Crivello beside a window
Nest08
OUR SERVICE COMMITMENT

We do the ML. You build the product.

Lower cost and p95 latency. Continuously maintained by our ML team.

≈10

Labeled references.

Once per agent/task.
They define correct labeling.

Not the full training set.
R$0

Upfront fees.

Training, evaluation
and deployment included.

≈24h

New frontier quality.

Match or beat it
on your agreed task.

Timing depends on data volume.
Nest’s service commitment. The approximate response window starts from new frontier model availability.
Nest09

Our ML team’s platform.

Our team uses it to train, evaluate, serve and improve your models.

Production traces

Start from real agent inputs and outputs.

PII removal

Remove sensitive fields before training.

Nest10
WHY THE ECONOMICS CAN WORK

Frontier-beating models.
Training runs under $100.

PredibaseLoRA Land · Predibase
<$8average / model

25 specialists ≥ GPT-4

Mistral 7B · task-specific evaluation

OpenPipeOpenPipe
~$80final training run

Email research: 96% vs. 90%

Qwen 14B specialist vs. OpenAI o3

Predibase, 2024 · OpenPipe ART, 2025 · OpenPipe RULER · Training-run costs exclude staffing, data work, experiments and serving.
Nest11

H100 rental prices fell 57%.

NVIDIAH100 SPOT-CONTRACT COMPOSITE INDEXUSD / GPU-HOUR
$0$2$4$6 $6.622H 2023$3.91Q4 2024$3.34Q2 2025$2.82JAN 2026 $0.62 202720282029JAN 2030
ObservedIllustrative projection
SemiAnalysis H100 index · Selected observations across providers and terms. Oct 2023 represents 2H 2023. Projection is Nest’s illustrative scenario.
Nest12

Customer support

1,245training examples

BC5CDR

5,228training examples

DBpedia

560,000training examples

CoNLL++

14,041training examples

GSM8K

7,473training examples

Magicoder

75,197training examples

WikiSQL

56,355training examples

Legal

17,000training examples

14,041 examples

Median training-set size across all 31 LoRA Land tasks.

LoRA Land, 2024 · Table 1 · Median of all 31 train splits. Not a minimum or performance guarantee.
Nest13

Synthetic data, at scale.

14,041 examples per training set (median) · 72% of fine-tuned models beat GPT-4 in LoRA Land.

API cost · USD (log scale)Time · minutes (estimated)Generate + label: 6k input + 1.75k output tokens/example
$4,002
≈275 min
2023
GPT-4
$46.34
≈47 min
2026
GPT-5.6 Luna
≈$0.24
≈25 min
2030 scenario
Illustrative projection
LoRA Land · Prices: OpenAI ’23 / ’26 · Throughput: Willows / AA · Independent scales; time estimated; 2030 projected.
Nest14

Small models. Serious capability.

MODELINTELLIGENCE INDEX v4.3TOTAL PARAMETERS
27B27B active
428B23B active
284B13B active
1.6T49B active
2.4T95B active
Artificial Analysis · Intelligence Index v4.3 · Sep 2026 · Total and active parameters shown separately. Parameter count ≠ inference cost; reasoning settings vary.
Nest15
Nest16
COMPETITIVE LANDSCAPE

Competitive positioning

ML work handled for you More is better
Customization More is better
Illustrative 0–10 positioning by Nest, not benchmark scores. Nest represents target scope; competitor scope varies by offering.
Nest17
Business model

You pay for SLM specialist inference.

Per call · illustrativeFrontier baseline $0.010

$0.001 per SLM specialist call

Serving cost

$0.0005 / call

Nest spread

$0.0005 / call

Your savings

$0.009 / call
Illustrative economics. Nest’s spread is before training and other delivery costs.
Nest18
THE PEOPLE BEHIND NEST

Built by an operator. Close to the problem.

Guilherme Martin

Guilherme
Martin

FOUNDER

Second-time founder.
Hands-on experience building custom models for real workflows.

PEOPLE IN OUR NETWORK
Thiago Bonini

Thiago Bonini

Founder, Vidia
Stanford MSx

Gustavo Gadotti

Gustavo Gadotti

Founder, Zipper
Acquired by CRM&BONUS

André Corrêa Santos

André Corrêa Santos

Machine Learning Engineer
Nubank

Nest19
ENTER · ILLUSTRATIVE ANNUAL SCENARIO
Current modelR$113.9M

R$52.4M

With specialists · per year

Current modelANNUAL BILL
With specialists
Illustrative savings · assumptions in notes.
Nest20
Market consolidation

Strategic buyers are acquiring the AI stack.

A new generation is still entering.

Selected acquisitions across the AI stack · Public filings and company announcements · Each card links to its source.
Nest21

If you needed
heart surgery,
which hospital
would you choose?

Nest22
Nest

Let’s make your
agents better.

Start with one agent.
Measure the difference.

EMAILguilherme@nest.technology
PHONE(11) 94205-9080
An interlocking block landscape in Nest’s neutral colors.
THANK YOU
Nest23