Asterion AI

Asterion AI Knowledge Platform

Asterion is the living knowledge base powering Cosmos Genesis. It stores physics concepts, architectural decisions, and simulation provenance — surfacing them through natural language and five context-adapted personas inside the studio.

Public reference. This page explains what Cosmos Genesis does and the science it's based on. Implementation details, such as exact algorithms, tuning values and infrastructure, are kept in our internal documentation.

What Asterion Knows — and What He Doesn't

A core design goal for Asterion is a 99%+ non-hallucinatory experience. This is enforced structurally — Asterion answers from Canon, not from LLM training data — and behaviourally: when something isn't in Canon, Asterion says so rather than guessing.

When Asterion knows

Answers are grounded in Canon items — every response carries the authority level, freshness class, and confidence score of the source evidence. You can see where the answer came from and how recently it was validated.

When Asterion doesn’t know

Asterion says so plainly. He does not guess, does not claim certainty beyond what evidence supports, and does not fill silence with invention. The response names what’s missing from Canon so you know exactly what the gap is.

When Asterion knows something adjacent

Asterion surfaces what Canon does contain that’s related, explains why it’s related, and points toward where the specific answer you want can actually be found — whether that’s a generation run, a validation report, or a domain that hasn’t been surveyed yet.

How the architecture enforces this

Hallucination prevention isn't a prompt instruction — it's a retrieval architecture. Asterion only answers from items that pass authority, freshness, and confidence gates. Items below threshold aren't surfaced, and the Q&A service is designed to detect empty retrieval results and respond with an explicit acknowledgment rather than a generated approximation.

  • →Retrieval-grounded: answers come from Canon, not LLM weights
  • →Authority gates: items below approved status are excluded from Q&A retrieval
  • →Freshness gates: expired items are excluded or flagged before surfacing
  • →Confidence propagation: source evidence confidence chains through to the answer
  • →Empty retrieval → explicit unknown: no answer generated when Canon has nothing

The Canon

The Canon is a governed, append-only knowledge store — the source of truth that all three AI identities and human reviewers read from and write to. Items enter as drafts and become active only after validation. Nothing is silently overwritten.

What Canon Stores

  • →Architectural decisions and their rationale
  • →Physics concepts and simulation methodology
  • →Process documentation and runbooks
  • →Simulation experiment results with provenance chains
  • →Project history and decision audit trail

Knowledge Kinds

  • →decision — architectural or product choices made
  • →concept — explanations of how something works
  • →process — workflows and runbooks
  • →architecture — system design documentation
  • →requirement — constraints and must-haves
  • →experiment — simulation result records
  • →configuration — environment and tuning values

Knowledge Lifecycle

Draft
↓
Active
↓
Superseded
↓
Archived

Items promoted from experiments enter as candidate_promoted_from_experiment and require human approval before becoming active. Supersession chains are tracked explicitly — no silent overwrites. All transitions are recorded in an immutable audit log.

Retrieval Architecture

Asterion answers questions by retrieving from Canon — not by generating from training data. The retrieval strategy prioritises governance over similarity: authority and freshness filter results before semantic ranking is applied.

Metadata-First Hybrid Search

Every query passes through five ranked signals before results are returned. An item that is semantically similar but stale or low-authority will rank below a fresher, more authoritative item — preventing outdated Canon entries from surfacing ahead of current ones.

Semantic scorecosine similarity between query and item embeddings
60%
Authority scorecanonical > approved > proposed > experimental > draft
20%
Freshness scorefresh (≤14d) → aging (≤30d) → stale (≤60d) → expired (>60d)
10%
Validation scorevalid > pending > conflicting > stale > failed
5%
Confidence scorepropagated from source evidence and auditor findings
5%

Semantic similarity (1024-dimension vector search) is applied only to the pre-filtered, authority-ranked candidate set — not to the full Canon.

The Auditor Corps

Six bounded specialist auditors maintain Canon health. Each has a single clear responsibility, runs on its own schedule, and publishes findings to the evidence ledger — never directly modifying active Canon. Human review is required before any auditor finding reaches active status.

Survey

Every 60 min

Discovers new knowledge from configured data sources and creates draft Canon entries for human review.

  • →Scans universe database (galaxies, star systems, planets)
  • →Scans codebase modules, API routes, and configuration
  • →Deduplicates against existing Canon before creating new items

Validity

Every 30 min

Compares canonical truth to observed reality — detects stale entries and supersession conflicts.

  • →Flags items not validated within 30 days
  • →Detects supersession conflicts (active items with broken supersession links)

Pattern

Daily 2 AM

Detects Canon structure anomalies and galaxy physics outliers. Auto-promotes findings with Measured (deterministic) confidence.

  • →Canon structure: orphaned items, unlinked clusters, repeated patterns
  • →Galaxy physics: virial ratio, spectral-type distribution, stellar mass outliers
  • →Promotes Measured findings to proposed Canon status; Estimated findings always go to human review

Simulation

On event

Validates client GPU experiment results and decides whether findings are eligible for Canon promotion.

  • →Evaluates experiment results: CONFIRMED / REJECTED / INCONCLUSIVE
  • →Promotes confirmed findings with Measured confidence to proposed Canon status; Estimated confidence always goes to human review
  • →Raises anomaly alerts when validated Canon items are contradicted by sensor data

Archive

Daily

Flags stale, superseded, and orphaned items for human review. Never auto-archives — all findings go to the review queue.

  • →Items with no validation for > 90 days
  • →Supersession chain breaks (A → B → C where B is missing or inactive)
  • →Dead galaxy references (items pointing to non-existent galaxy IDs)

Codebase

Nightly (dev only)

Validates that code and documentation stay aligned — detects drift between architectural intent and implementation.

  • →Module and API docstring accuracy (large language model (LLM) evaluation)
  • →Broken file path references in documentation
  • →Architecture drift: planning docs vs. deployed CDK stacks
  • →Stale TODO/FIXME markers (> 30 days old)
  • →Single responsibility violations and cyclomatic complexity > 10

Simulation Provenance

Without provenance, a simulation result is just a number. You cannot reproduce it independently, submit it for peer review, audit it for a grant claim, or trust it for a mission-critical decision without knowing exactly what was running when it was produced. Provenance chains solve this: every output is linked to the precise inputs, physics model version, and generation parameters that produced it — making any result independently reproducible from the chain alone.

Without provenance

Simulation outputs cannot be reproduced by an independent party. There is no basis for peer review, no audit trail for grant reporting, and no way to establish confidence in a trajectory or burn decision beyond taking the platform’s word for it.

With a provenance chain

Any result is independently reproducible from the chain alone — no internal state access required. Peer reviewers, grant auditors, and mission planners all have the same verifiable reconstruction path. This is what distinguishes a scientific simulation platform from a calculator.

Connection to Canon promotion

The Simulation Auditor can only confidently promote an experiment result to proposed Canon status when the full provenance chain is intact. A broken or incomplete chain blocks promotion — keeping unverifiable data out of the knowledge base.

Event-Driven Provenance Tracking

Outputs of simulations are verified through event-driven deviation checking. Every simulation result carries a provenance chain linking it to the source inputs, physics model version, generation parameters, and the deterministic random seed that produced it. Confirmed results are promoted to proposed Canon status; results that contradict active Canon items raise anomaly alerts for human review.

  • TraceabilityEvery output linked to source inputs, model version, and generation parameters
  • Deterministic seedThe physics pipeline uses hierarchical deterministic seeding — same seed + same parameters reproduces an identical universe. The seed is a required part of a complete provenance chain.
  • VerificationProvenance chains support independent reproduction without internal state access — only the chain is needed
  • Deviation detectionEvent-driven N-body output comparison against expected trajectories
  • Grant & peer reviewMeets NSF/NASA reproducibility requirements — the chain is the audit artifact
  • Mission safetyIf a simulation informed a trajectory or burn decision, the chain proves what was running when

What Confirmed Runs Enable

The confirmed run corpus isn't an archive — it's an active input to five downstream processes. A provenance chain is what makes a run eligible to participate in all of them.

Anomaly baselineConfirmed Canon items become the reference signal. Future sensor readings that contradict a validated item trigger an AnomalyAlert — without confirmed runs, there is nothing to contradict.
Physics calibrationThe Pattern Auditor aggregates virial check results into a rolling 30-day summary (mean Q ± std dev, % confirmed). If Q consistently drifts from 0.5, the N-body initialization is mis-tuned. The corpus makes that visible.
Prediction trackingPredicted orbital events become pending-observation Canon items. The Simulation Auditor sweeps daily for expired prediction windows — runs feed the predictions; Canon tracks whether reality confirms them.
Q&A groundingEmpirical experiment results are retrievable. “What’s the typical virial ratio across generated galaxies this month?” is answered from measured Canon data, not training weights.
Emergent discoveryWhen users independently confirm the same unexpected physics behavior in generated systems, that’s a finding that didn’t exist in the inputs. Canon is where it becomes a citable result.

Natural Language Interface

Asterion answers natural language questions by consulting Canon, then optionally calling tools to retrieve live simulation state. Questions about objects that exist in your workspace are answered from measured data — not inferred from training.

Universe Tools

  • →Create generation jobs for new galaxies or star systems
  • →Predict orbital events (conjunctions, oppositions, transit windows)
  • →Design gravitational capture scenarios
  • →Add interstellar visitors to existing systems
  • →Fetch validation reports for generated objects

Experiment Tools

  • →Save and load simulation experiment records
  • →List experiments by galaxy or user
  • →Generate system specifications from experiment results
  • →Submit GPU client results for Simulation Auditor review

Workspace-Aware Context

Asterion doesn't just answer questions about abstract physics and architecture. He knows what you're working on right now — the scene you have loaded, the galaxies you've generated, the simulations currently running. Every answer is scoped to your universe, not a generic knowledge base.

Example — live workspace response

“What is before us now: Your scene contains a galaxy called ‘New Stellar Cluster’ — currently empty of star systems. The Auditors have brought me detailed measurements of stellar populations across multiple spectral types: from the rare, massive O-type stars (averaging 20.7 solar masses) down through the abundant M-type dwarfs (averaging 0.2 solar masses). These are verified records from the cosmological production authority.”

Asterion knows your workspace

The current scene, loaded galaxies, running simulations, and generated objects are all part of the context Asterion answers from. Asking “what stars are in this system?” returns measured data from your generation run — not textbook averages.

Answers are scoped to your universe

Canon is populated by Auditors surveying your actual generated universe. When you ask about stellar populations, virial ratios, or object counts, Asterion is reporting on measurements taken from your data — not returning generic astrophysics definitions.

This is the distinction between a knowledge platform and a chatbot

A chatbot answers from training weights. Asterion answers from Canon items that were created by Auditors analyzing your generated universe — grounded in your specific objects, parameters, and simulation results. The difference is measurable: every answer carries the source, freshness class, and confidence of the evidence behind it.

What workspace context includes

  • →Active scene — the galaxy or system currently loaded in the studio
  • →Generated objects — galaxies, star systems, planets, and celestial bodies from your generation runs
  • →Running simulations — N-body experiments and mission trajectories in progress
  • →Validated measurements — stellar population surveys, virial ratios, and spectral distributions from the Auditor Corps
  • →Generation parameters — the seeds, physics model versions, and configuration values that produced your universe

Edge & Air-Gap Deployment

Cosmos Genesis is designed to operate in environments where cloud connectivity is unavailable, unreliable, or prohibited. The Canon export architecture was built with this in mind from the start — targeting space operations, emergency response, and secure facility deployments where sneakernet is the only available data channel.

Signed Canon Bundles

  • →Canon exports are Ed25519-signed bundles — integrity is verifiable without a network connection
  • →Point-in-time named snapshots enable version-pinned offline deployments
  • →Organization-scoped exports carry full provenance chains for every item
  • →Bundle verification requires no cloud call — signature check is fully local

Target Environments

  • →Deep space missions — intermittent or absent uplink windows
  • →Emergency response — field deployments with no reliable infrastructure
  • →Secure facilities — air-gapped networks where cloud egress is prohibited
  • →Remote operations — areas where connectivity is available only in windows

On-Device / Cloud Latency Indicator

Every Asterion response shows the time spent on-device versus in the cloud — for example, 0.0s on-device · 5.0s cloud. Canon retrieval, validation, and signing are fully local and complete in milliseconds. Cloud latency reflects Bedrock inference time. As local LLM support expands, the cloud figure drops — the indicator makes that progress concrete for every response.

  • →Canon retrieval, signing, and validation: fully local — typically 0.0s
  • →Embedding generation: cloud (Bedrock Titan) today, local model planned
  • →Q&A inference: cloud (Bedrock Claude) today, local LLM runtime in development
  • →Indicator is per-response — operators see the split on every interaction

Active development direction. Canon export, signing, retrieval, and validation are production-ready today. Local embedding inference and offline Q&A are in active development — the cloud/local indicator tracks progress toward a fully air-gapped deployment.

Personas

Asterion is designed to surface Canon through five personas — each adapted to a distinct user context. The underlying knowledge base is shared; the delivery would adapt to who is asking.

Mission Controller

Operational context. Focused on mission parameters, orbital mechanics, and real-time simulation state.

Lab Assistant

Technical depth. Answers physics and engineering questions with precision — appropriate for researchers and engineers.

Educator

Accessible explanations. Prioritizes clarity and analogy over technical detail — appropriate for students and general audiences.

Guide

Discovery orientation. Helps users explore the generated universe — surfaces interesting objects, phenomena, and context.

Game Master

Narrative framing. Translates physics data into story-relevant context for game developers and interactive experience designers.

Planned, not yet implemented. The five personas above describe the intended delivery layer. Today, every query is answered through Asterion's single unified interface — persona-specific framing has not been built yet.

AI Engineering Pipeline

How Asterion integrates into the AI-human hybrid development methodology — the three AI identities, Codebase Auditor review, and human architectural oversight.

Read →
← Back to docs