Starling Labs · Applied Research

Your platform AI is ready. Your data is not.
We fix it inside your own tenant. Nothing leaves.

We train small industry models, ground them in knowledge graphs, and run them inside your own tenant. The output is clean data to maximize your existing AI investments.

95%+
Verified remediation, all five models
0.90
Confidence gate to auto-apply
4B to 14B
Runs in your tenant
Weeks
To production
Starling Sovereign

Why Sovereign matters.

Every answer this suite produces is computed inside your own tenant. That is what Sovereign means, and it is not a deployment detail. Your data never crosses a boundary you don't control: the models are small enough to live where the data lives, the knowledge graph runs beside your systems of record, and every action a model or agent takes is written to a tamper-evident ledger your auditors can verify.

Most enterprise AI asks you to trade custody for capability. Sovereign exists so you never make that trade. For a regulated enterprise, that is the difference between a pilot and production.

Your tenant, your data, your proof.

Starling Sovereign

Remediation vs. Cleaning

Cleaning fixes the value. Remediation fixes the value and keeps the record of it: what the defect was, how confident the system was, who or what approved the change, and what it looked like before.

Cleaning gives you a better dataset. Remediation gives you a better dataset plus the evidence that proves how it got that way.

Sovereign doesn't just clean. Sovereign remediates.

Overview

Watch how it works.

No autoplay. Each video loads only when you press play.
Announcements

From the lab.

STARLING RESEARCHRepair youcan defend.Source-grounded remediation: every correctionverified against the record of truth,or honestly escalated.RECORDSunnyval​leSOURCE OF TRUTHSunnyvaleverified restore · conf 0.97
Research2026 · 09

Repair you can defend.

Source-grounded remediation: every correction verified against the record of truth, or honestly escalated.

Research note · 012026 · 09

Repair you can defend.

Source-grounded remediation: every correction verified against the record of truth, or honestly escalated.

When a Starling model finds a defect, it does not guess at a fix. It reaches back to your source of truth, the baseline your CMDB, ledger, or registry already holds, and restores the verified value. Each restore carries the evidence with it: the field, the before and after, the source it was verified against, and the confidence behind the change.

When no verified value exists, the model says so. The record is escalated to review with a plain-language rationale instead of a silent guess. That is the whole contract: every repair defensible, every uncertainty honest.

DETERMINISTICLEARNED JUDGMENTSTARLING RESEARCHRules where rules win.Models where judgment is needed.
Research2026 · 09

Rules where rules win. Models where judgment is needed.

Deterministic checks handle what rules do best; learned judgment takes over where rules run out.

Research note · 022026 · 09

Rules where rules win. Models where judgment is needed.

Deterministic checks handle what rules do best; learned judgment takes over where rules run out.

Format checks, code lists, referential integrity: rules are perfect for these, and Starling runs them as rules. They are fast, free, and never wrong about what they can see.

But rules run out. A rule cannot tell a rebranded vendor from a typo, or a plausible-but-stale value from a current one. That is where the learned layer takes over, applying judgment the way a domain expert would, with its confidence stated on every call. Each layer does only what it is best at, and the seam between them is explicit.

CLEAN RECORDtruth establishedDEFECT INanswer knownREMEDIATEmodel proposesVERIFY== known truth?only verified answers trainSTARLING RESEARCHEvery answerknown.Ground-truth verified training: models learnonly from remediations provencorrect by construction.
Research2026 · 09

Every answer known.

Ground-truth verified training: models learn only from remediations proven correct by construction.

Research note · 032026 · 09

Every answer known.

Ground-truth verified training: models learn only from remediations proven correct by construction.

Our training loop starts from a clean record whose truth is established. We inject a defect, so the right answer is known before the model ever sees the problem. The model proposes a remediation, and the proposal is verified against that known truth.

Only verified answers train. A repair that cannot be proven correct never enters the training set, which is why the models’ accuracy figures are measured, not estimated. Correct by construction, verified before learned.

01 · Thesis

The gains come from elsewhere.

The next gain is not a bigger model. It's architecture, training, and agents that act.

Frontier labs race to scale general models. Enterprises need the opposite. Most enterprise value is stranded. Not for want of a larger model, but for want of models tuned to the task, architectures that stay cheap at length, and agents that can act and be trusted to.

We work that frontier across five lines at once: specialized models, post-transformer architectures, agentic systems, knowledge graphs, and the governance that holds them together. It ships as one suite: Starling Sovereign. And we ship it.

02 · What we build

Five research lines.

Distinct efforts, one suite: Starling Sovereign. Built by the team that measures it.

R · 01

Specialized models

Five industry models: Starling Health, IT, Retail, Property, and Finance. Each is distilled from a 120B teacher with ground-truth verification, reaches 95%+ verified remediation, and runs in your tenant at 4B, 8B, or 14B.

120B teacherground-truth distillation4B · 8B · 14Bmodel cards →
R · 02

Beyond attention

State-space and hybrid attention architectures, and mixture-of-experts. Linear-time inference and long context where it pays.

SSMhybrid attnMoEO(n)
R · 03

Agentic harnesses & swarms

A planner decomposes goals to specialist agents over an MCP tool bus, with guardrails, sandboxed execution, and full run tracing.

planner / specialistsMCPguardrails
R · 04

Graph-native intelligence

A knowledge graph for semantic resolution and federated merge. Graph-RAG, the substrate agents reason over, not a rule per variant.

knowledge graphGraph-RAGentity resolution
R · 05

Provenance & governance

Across every line above, a SHA-256 chained, canonical (RFC 8785) audit trail of what each model and agent did. Identity, inputs, and attribution, with multi-year retention. The accountability that lets autonomy run in a regulated industry.

SHA-256 chainRFC 8785 / JCSmodel cards →runtime governance
03 · Results

Measured, not promised.

95%+
Verified remediation across all five industry models.
99.6%
Starling Property, the strongest in the family.
100%
Schema-valid output from every model, ready for the audit ledger.
Measured on held-out public-data evaluations never seen in training. We validate against your data before standing behind a number.
Read the model cards →
04 · Approach

How we work.

A · 01

Empirical by default

Every model benchmarked on held-out data never seen in training before it ships. Claims are numbers, not demos.

A · 02

Architecture-agnostic

Transformer, state-space, hybrid, or MoE. The architecture the task rewards, not the one in fashion.

A · 03

Agents under governance

Harness, guardrails, and tracing. Autonomy and accountability ship together.

A · 04

Research that ships

The team that trains it owns it in production. No handoff, no lab-to-field gap.

Starling Sovereign · Model Cards

Every model, on the record.

Five production models, one contract. Each remediates the records its industry runs on, inside your own tenant. All metrics are measured on held-out public-data evaluations never seen in training. Every model produces 100% schema-valid output and was trained via ground-truth-verified distillation from a 120B teacher.

Verified-remediation rate = detection on source-covered fields.
Unsupervised error rate = auto-applied false edits at the 0.90 confidence threshold.
MC · 01 · Starling Health

Starling Health

Clinical-grade data hygiene for healthcare IT.

Starling Health remediates the records that healthcare runs on: provider registries, credentialing data, clinical application inventories, and device fleets. Trained on millions of defects injected into real public registry data (NPPES-class), it detects stale taxonomy assignments, corrupted identifiers, cross-field conflicts, and exposed PHI, then either restores the verified value from your source-of-truth baseline or masks and escalates. It never invents a value it cannot verify.

DriftCorruptionConflictCompliance · PHI/PIIIntegrity
95.9%
Verified-remediation rate
0.3%
Unsupervised error rate
Auto-applies only what it can defend

Intended use

  • Provider registries and credentialing data
  • Clinical application inventories and device fleets
  • PHI: detect-and-mask with full audit rationale on every correction

Training and evaluation

  • Trained on millions of injected defects in real public registry data (NPPES-class)
  • Held-out public-data evaluation, never seen in training
  • Ground-truth-verified distillation from a 120B teacher
MC · 02 · Starling IT

Starling IT

A self-healing layer for your CMDB.

Starling IT keeps configuration data honest: servers, applications, device firmware, vulnerability references. Trained against live CVE/CPE vocabularies, it recognizes canonical infrastructure identifiers on sight, including CPE URIs, CVE ids, and CVSS scores, and repairs drifted versions, broken references, and placeholder rot without ever reformatting a value that is already canonical.

DriftCorruptionConflictIntegrity
96.6%
Verified-remediation rate
0.5%
Unsupervised error rate

Intended use

  • Servers, applications, device firmware, vulnerability references
  • Continuous reconciliation against your CMDB baseline

Training and evaluation

  • Trained against live CVE/CPE vocabularies
  • Native fluency in CVE/CPE/CVSS formats: zero false "fixes" of canonical identifiers
  • Held-out evaluation; distilled from a 120B teacher
MC · 03 · Starling Retail

Starling Retail

Catalog quality at commerce scale.

Product data is the dirtiest data in the enterprise: crowdsourced attributes, multilingual names, inconsistent taxonomies. Starling Retail was trained on genuinely messy public catalog data, not sanitized samples, so it knows the difference between an unusual-but-valid product name and a real defect. It normalizes categories, repairs corrupted attributes, and reconciles conflicting listings while leaving merchandising voice intact.

DriftCorruptionConflictIntegrity
96.9%
Verified-remediation rate
0.0%
Unsupervised error rate
On evaluation: everything uncertain routes to review

Intended use

  • High-volume catalog ingestion pipelines
  • Category normalization, attribute repair, listing reconciliation
  • Multilingual-aware; tag-taxonomy fluent

Training and evaluation

  • Trained on genuinely messy public catalog data, not sanitized samples
  • Held-out evaluation; distilled from a 120B teacher
MC · 04 · Starling Property

Starling Property

Trustworthy records for real estate and property data.

From tax-lot rolls to listing feeds, property data mixes free-text addresses, coded classifications, and high-stakes assessed values. Starling Property, trained on municipal-scale public property records (NYC PLUTO-class), validates zoning and building-class codes, repairs malformed addresses, and reconciles assessment figures against authoritative baselines, with every change carrying its rationale and confidence.

DriftCorruptionConflictIntegrity
99.6%
Verified-remediation rate
The strongest in the family
0.5%
Unsupervised error rate

Intended use

  • Assessor, brokerage, and prop-tech data pipelines
  • Zoning and building-class code validation; address normalization
  • Assessment reconciliation against authoritative baselines

Training and evaluation

  • Trained on municipal-scale public property records (NYC PLUTO-class)
  • Held-out evaluation; distilled from a 120B teacher
MC · 05 · Starling Finance

Starling Finance

Audit-ready data quality for financial records.

Financial data carries the highest cost of a silent error: vendor masters, general-ledger references, cost-center hierarchies, and payment records where a drifted code or conflicting entry becomes a reconciliation problem or an audit finding. Starling Finance validates coded classifications, repairs corrupted references, and reconciles conflicting entries against your authoritative baseline, with every change carrying its rationale and confidence for the audit ledger.

DriftCorruptionConflictComplianceIntegrity
95.22%
Verified-remediation rate
0.90
Confidence threshold
Everything uncertain routes to review

Intended use

  • Vendor masters, general-ledger references, cost-center hierarchies
  • Payment and reconciliation records
  • Every correction audit-ready: rationale and confidence on each change

Training and evaluation

  • Trained on defects injected into public financial-registry data
  • Held-out public-data evaluation, never seen in training
  • Ground-truth-verified distillation from a 120B teacher
Shared platform

How every Starling model works.

Source-grounded remediation

Every model cross-checks records against your source-of-truth baseline (CMDB, ledger, registry) before proposing any change.

Confidence-gated autonomy

Corrections at or above 0.90 confidence auto-apply; everything else lands in the review queue with a plain-language rationale.

Full provenance

Every correction ships as structured JSON: field, before/after, severity, defect class, policy, rationale, confidence, ready for the audit ledger.

Runs where your data lives

Same contract at every tier, scaled accuracy. On-prem including Apple Silicon.

4B
Edge
Lightest footprint
8B
Standard
The production default
14B
Enterprise
Highest accuracy

Talk to the lab.

Enterprises, partners, researchers. One conversation, no funnel.

Book a call

Tell us where to start.

Share your area of interest and contact details. Your note goes straight to the team.

Sends directly to the team. CAPTCHA protected.