Open-Source · Edge AI Research05

INTELLIGENCE ISN'T
ABOUT SCALE. IT'S
ABOUT precision.

Pluto AI Labs is an open-source research lab building SOTA sub-7B language models, reasoning engines, and MLOps tooling for edge deployment. We prove elite AI can run on consumer hardware.

ATLAS-CODER (0.5B)
36.6%EvalPlus strict
Apollo-VL-Edge-3B
#2Sub-5B VLM
Cumulative downloads
9,000+across model releases
Researcher working with edge-deployed AI models at Pluto AI Labs
MBPP+ score
43.9%
Atlas-Coder-2 · 0.5B
○ Hugging Face○ PyPI○ GitHub○ EvalPlus○ Zenodo / CERN○ IRJMETS
Compact edge-AI hardware form factor
§ About Us

THE
mission.

Knowledge DistillationHigh-Signal DataEdge MLOps

The AI industry is obsessed with scaling up — building massive, bloated models that require data center GPUs to run. At Pluto AI Labs, we are obsessed with scaling down.

We believe elite reasoning, coding logic, and visual understanding shouldn't require a $30,000 GPU cluster. By rigorously engineering high-signal datasets, applying Knowledge Distillation, and building fault-tolerant MLOps pipelines, we compress frontier-level intelligence into 1B–7B parameter models that run locally, offline, on consumer laptops.

WHAT WE DO

Three layers of one ecosystem: the models, the data that trains them, and the tooling that keeps them honest.

The Brains

Models

We don't just fine-tune; we distill frontier behaviors into efficient edge packages.

  • Apollo-VL-Edge-3B (Flagship VLM) — a 3B Vision-Language Model fine-tuned through a 210-hour 2x T4 DDP QLoRA pipeline. It beats Alibaba's Qwen2.5-VL-3B on ChartQA (78.6%) and AI2D (77.9%). #2 Sub-5B VLM. Runs on 8GB VRAM consumer GPUs at 36 tok/s. 450+ downloads in 48 hours.
  • Atlas Collection (Coding) — SOTA sub-1B coding models. Atlas-Coder-2-0.5B hits 36.6% HumanEval+ and 43.9% MBPP+ (EvalPlus strict), beating Qwen2.5-Coder and DeepSeek-Coder.
  • Pluto Collection (General) — general-purpose edge LLMs for multi-step reasoning. Pluto-Genesis-0.6B ranks #4 globally among sub-1B baselines.
Browse on Hugging Face
The Fuel

Datasets

High-signal density beats raw scale. We open-source our rigorously cleaned training data.

  • Atlas-Frontier-Model-Traces (15K rows) — distillation traces from Kimi-K3, GPT-5.6 and Fable-5, with 56% of lazy teacher-model failures aggressively filtered out.
  • Orbit-200K (200K rows) — a universal ChatML dataset (math + code + general instruction) engineered for zero-bloat reasoning.
  • Atlas-Coder-50K — 50,000 execution-verified Python samples for specialized coding fine-tuning.
  • Apollo-VL-Massive-Dataset (161,562 rows) — a multimodal dataset engineered to train VLMs to perform structured visual reasoning inside <think> tags before answering. Built for knowledge distillation, behavioral cloning, OCR, chart understanding and multimodal instruction tuning of compact VLMs. Crossed 30+ downloads in the first 12 hours of release.
View datasets
The Tooling

Infrastructure

We build the dev tools we wish existed for local AI development.

  • llm-diff (published on PyPI) — the git diff for LLM behavior. A CLI that probes two models with behavioral tests and prints a colored diff table of instruction fidelity, verbosity and reasoning consistency.
  • Built for CI/CD pipelines, so silent model regressions get caught before deployment.
Install pluto-llm-diff
§ Impact & MetricsAs of August 2026
New releaseApollo-VL-Edge-3B — 450+ downloads in the first 48 hours
9,000+
Cumulative model downloads
4
Open-source datasets published
#top 5
Sub-1B coding model (EvalPlus)
#2
Sub-5B Vision-Language Model
1
PyPI package — pluto-llm-diff
2
Peer-reviewed research papers

THE ROADMAP —
WHAT'S next.

The Pluto AI Labs Roadmap: proving that data beats scale with Athena 1.0, a 4B model engineered to match 9B-class vision-language intelligence.
Athena 1.0 vision-language model running locally on consumer hardware
In progress
Next Frontier · Q4 2026

Athena 1.0 — 9B-Class Intelligence in a 4B Model

Apollo-VL proved that a 3B model can see and reason. Athena proves it doesn't need 9 billion parameters to do it at 9B level.

We are forging a 4B vision-language model on a 250,000-row precision dataset — the Athena 1.0 Brain — curated from ChartQA, Docmatix, and The Cauldron's deepest reasoning families. Every row is decontaminated, deduplicated, and identity-locked before it ever touches the weights. The mission: match previous-gen 9B models on document, chart, and OCR intelligence — on a laptop, offline — and prove once and for all that data beats scale. Built on Qwen3-VL-4B. Trained on full fine-tuning across chained premium GPU sessions with crash-proof checkpoint recovery. Ships with a locked identity ('I'm Athena 1.0'), quantized for edge, benchmarked in the open. Precision over parameters. That's the Pluto thesis — Athena is its flagship proof.

Active roadmapFollow releases

RESEARCH &
publications.

We back our engineering with academic rigor — published, reproducible, and open for critique.

The Distillation Bottleneck

LinkedIn / HF Blog

A technical study on how raw distillation datasets contain up to 56% lazy failures, and the data-centric MLOps pipeline required to filter them for high-signal training.

Pluto-Genesis: Reliable QLoRA Fine-Tuning on Ephemeral Compute

Zenodo / CERN

Engineering a 3-layer cross-session checkpoint recovery system to bypass 12-hour hardware limits.

SENTIENT AUDIT

IRJMETS

An AI-driven multi-pass security analysis platform for smart contract vulnerability detection.

Run frontier-level intelligence locally with Pluto AI Labs models

Every model, dataset, and tool we ship is open weight and open source — offline-capable, consumer-hardware first, no cluster required.