PUBLIC RESEARCH SIGNAL

Pushing the limits of agent capability.

Gaia Research is the open laboratory for evidence-first agent work. We observe, benchmark, verify, and publish the frontier — every claim links back to its receipts.

Milim, Gaia Research's Chief Capability Scout, standing in a laboratory hoodie.

MILIM · CHIEF CAPABILITY SCOUT

VERIFIED EVIDENCETWO LABS, ZERO SIGN-UPSBUILDING IN PUBLIC

FROM THE BLOG

A little field research.

Short, human-readable notes from the Gaia lab—where the work gets a little strange before it gets useful.

All blog posts

PULL A LEVER, SEE WHAT SPARKS

Labs & Games.

Little browser toys we build to poke at agent capability — no login, no download, nothing leaves your machine. Go ahead, boss. Break something.

HELL HEAVEN INDEX · ARC I VERIFIED

Stop installing skills.
Start summoning them.

Marketplaces make you install skills forever — bloat you never asked for, pinned to every repo. Skill Zero is the launcher that starts a harness with zero ambient skill debt. Heaven and Hell are the summon directions on the HH axis: Heaven converges toward the right few skills; Hell explores the evidenced world.

New to Gaia? Start with the three names — Gaia Skill Tree, Gaia Research, and Gaia Skill Heaven (housing Skill Zero and Skill Hell). The HH Index is in the works.

☁ HEAVEN · AXIS WIP

Converge on the right few.

Heaven is the curated summon direction: it pulls the evidenced pool toward the small set a design session actually needs. Skill Zero gives it a clean launch when useful; Heaven decides what deserves to come back.

🔥 HELL · GATED

Full gas, autopilot.

Summon every good skill in the evidenced world for autonomous fleets and long loops — bounded by the rung you picked. Unlocks only when the registry’s trust-coverage clears a measured gate. Ludicrous mode ships with a seatbelt.

Skill Zero is the usable launcher prototype. The Hell Heaven (HH) Index — a per-skill hellHeaven stamp over Heaven/Hell polarity — is the research now in the works (routing currently falls back to relevance ranking). Read the benchmark method & receipts → · Vision ↗ · Mission ↗.

Skills you can install today.

Local-first Claude Code / Cursor / Windsurf skills. One line to install with npx skills — they run against your files and never upload their contents.

context-diet

WIP EXPERIMENTAL

Measure and compact an oversized agent-context file under the harness limit without dropping a single rule.

npx skills install gaia-research/skill-context-diet

cost

ACT ACTIVE

Multi-harness token-usage cost reporter for pi, Claude Code, Codex, and opencode session logs.

npx skills install gaia-research/skill-cost

fuse

ACT ACTIVE

Compose two installed agent skills into one unified SKILL.md — the composition engine behind the Skill Tree.

npx skills install gaia-research/skill-fuse

image-tuner

ACT ACTIVE

Interactive bounding-box image positioning, zoom, and framing tuner with live coordinate and CSS export.

npx skills install gaia-research/skill-image-tuner

Claims deserve a trail.

Explore research →
Selected public research ledger entries
Research itemTypeStatusEvidence note
Context DietLABWIP EXPERIMENTALLocal token-budget estimator; comparative benchmark results pending review.
The Compounding Cost of CI FailuresPOSTMORTEMVRF VERIFIEDPostmortem of Epic #780 introducing CI Churn as a first-class cost metric.
Agent Cost ReportingRESEARCH PLANPRP PROPOSEDProposed study of the gap between agent estimates, rate-card totals, and invoices.
Gaia MCPPRODUCTACT ACTIVEModel Context Protocol server exposing the Gaia Skill Tree to Claude Code, Codex, and Cursor.
Scout Fan-OutBENCHMARKVRF VERIFIEDEmpirical study of parallel cheap-scout fan-out vs single mid-tier scouts across 360 runs: Pareto frontier, prompt-cache amplification, and flake rate reduction.
HH BenchmarkBENCHMARKVRF VERIFIEDDrug-trial method for scoring skills by marginal efficacy against model baselines. Arc I verified: baseline floor pricing, census, and machine-gated claim index.

RESEARCH SIGNAL / PERMANENT RECORD

Let the research take root.

Gaia Research tests the frontier in public. Gaia Skill Tree keeps the attributable, evidence-backed record of what agents can actually do.

FROM
Open research, methods, and experiments.
TO
A public registry of demonstrated agent skills.
Enter Gaia Skill Tree (opens in a new tab)
Milim rests beneath the roots of a colossal luminous golden skill tree as petals fall through its branches.
A quiet place beneath the Skill Tree, where open research becomes a permanent record.

A field note from Milim, Chief Capability Scout.

Make the science
louder than the hype.

  1. 01 Share the method. Let others inspect the work.
  2. 02 Verify before you flex. Evidence beats vibes.
  3. 03 Build safe, robust agents. Then publish the receipts.
Explore the Skill Tree