Meet Starling Labs.
What the lab is, who it serves, and why the fix belongs inside your own tenant.
We train small industry models, ground them in knowledge graphs, and run them inside your own tenant. The output is clean data to maximize your existing AI investments.
Every answer this suite produces is computed inside your own tenant. That is what Sovereign means, and it is not a deployment detail. Your data never crosses a boundary you don't control: the models are small enough to live where the data lives, the knowledge graph runs beside your systems of record, and every action a model or agent takes is written to a tamper-evident ledger your auditors can verify.
Most enterprise AI asks you to trade custody for capability. Sovereign exists so you never make that trade. For a regulated enterprise, that is the difference between a pilot and production.
Your tenant, your data, your proof.
Cleaning fixes the value. Remediation fixes the value and keeps the record of it: what the defect was, how confident the system was, who or what approved the change, and what it looked like before.
Cleaning gives you a better dataset. Remediation gives you a better dataset plus the evidence that proves how it got that way.
Sovereign doesn't just clean. Sovereign remediates.
What the lab is, who it serves, and why the fix belongs inside your own tenant.
Architecture, training, knowledge graphs, and governance, for the evaluators.
The next gain is not a bigger model. It's architecture, training, and agents that act.
Frontier labs race to scale general models. Enterprises need the opposite. Most enterprise value is stranded. Not for want of a larger model, but for want of models tuned to the task, architectures that stay cheap at length, and agents that can act and be trusted to.
We work that frontier across five lines at once: specialized models, post-transformer architectures, agentic systems, knowledge graphs, and the governance that holds them together. It ships as one suite: Starling Sovereign. And we ship it.
Distinct efforts, one suite: Starling Sovereign. Built by the team that measures it.
Five industry models: Starling Health, IT, Retail, Property, and Finance. Each is distilled from a 120B teacher with ground-truth verification, reaches 95%+ verified remediation, and runs in your tenant at 4B, 8B, or 14B.
State-space and hybrid attention architectures, and mixture-of-experts. Linear-time inference and long context where it pays.
A planner decomposes goals to specialist agents over an MCP tool bus, with guardrails, sandboxed execution, and full run tracing.
A knowledge graph for semantic resolution and federated merge. Graph-RAG, the substrate agents reason over, not a rule per variant.
Across every line above, a SHA-256 chained, canonical (RFC 8785) audit trail of what each model and agent did. Identity, inputs, and attribution, with multi-year retention. The accountability that lets autonomy run in a regulated industry.
Every model benchmarked on held-out data never seen in training before it ships. Claims are numbers, not demos.
Transformer, state-space, hybrid, or MoE. The architecture the task rewards, not the one in fashion.
Harness, guardrails, and tracing. Autonomy and accountability ship together.
The team that trains it owns it in production. No handoff, no lab-to-field gap.
Five production models, one contract. Each remediates the records its industry runs on, inside your own tenant. All metrics are measured on held-out public-data evaluations never seen in training. Every model produces 100% schema-valid output and was trained via ground-truth-verified distillation from a 120B teacher.
Clinical-grade data hygiene for healthcare IT.
Starling Health remediates the records that healthcare runs on: provider registries, credentialing data, clinical application inventories, and device fleets. Trained on millions of defects injected into real public registry data (NPPES-class), it detects stale taxonomy assignments, corrupted identifiers, cross-field conflicts, and exposed PHI, then either restores the verified value from your source-of-truth baseline or masks and escalates. It never invents a value it cannot verify.
A self-healing layer for your CMDB.
Starling IT keeps configuration data honest: servers, applications, device firmware, vulnerability references. Trained against live CVE/CPE vocabularies, it recognizes canonical infrastructure identifiers on sight, including CPE URIs, CVE ids, and CVSS scores, and repairs drifted versions, broken references, and placeholder rot without ever reformatting a value that is already canonical.
Catalog quality at commerce scale.
Product data is the dirtiest data in the enterprise: crowdsourced attributes, multilingual names, inconsistent taxonomies. Starling Retail was trained on genuinely messy public catalog data, not sanitized samples, so it knows the difference between an unusual-but-valid product name and a real defect. It normalizes categories, repairs corrupted attributes, and reconciles conflicting listings while leaving merchandising voice intact.
Trustworthy records for real estate and property data.
From tax-lot rolls to listing feeds, property data mixes free-text addresses, coded classifications, and high-stakes assessed values. Starling Property, trained on municipal-scale public property records (NYC PLUTO-class), validates zoning and building-class codes, repairs malformed addresses, and reconciles assessment figures against authoritative baselines, with every change carrying its rationale and confidence.
Audit-ready data quality for financial records.
Financial data carries the highest cost of a silent error: vendor masters, general-ledger references, cost-center hierarchies, and payment records where a drifted code or conflicting entry becomes a reconciliation problem or an audit finding. Starling Finance validates coded classifications, repairs corrupted references, and reconciles conflicting entries against your authoritative baseline, with every change carrying its rationale and confidence for the audit ledger.
Every model cross-checks records against your source-of-truth baseline (CMDB, ledger, registry) before proposing any change.
Corrections at or above 0.90 confidence auto-apply; everything else lands in the review queue with a plain-language rationale.
Every correction ships as structured JSON: field, before/after, severity, defect class, policy, rationale, confidence, ready for the audit ledger.
Same contract at every tier, scaled accuracy. On-prem including Apple Silicon.
Enterprises, partners, researchers. One conversation, no funnel.
Share your area of interest and contact details. Your note goes straight to the team.