Asterion AI
Asterion AI Knowledge Platform
Asterion is the living knowledge base powering Cosmos Genesis. It stores physics concepts, architectural decisions, and simulation provenance — surfacing them through natural language and five context-adapted personas inside the studio.
Public reference. This page explains what Cosmos Genesis does and the science it's based on. Implementation details, such as exact algorithms, tuning values and infrastructure, are kept in our internal documentation.
What Asterion Knows — and What He Doesn't
A core design goal for Asterion is a 99%+ non-hallucinatory experience. This is enforced structurally — Asterion answers from Canon, not from LLM training data — and behaviourally: when something isn't in Canon, Asterion says so rather than guessing.
When Asterion knows
Answers are grounded in Canon items — every response carries the authority level, freshness class, and confidence score of the source evidence. You can see where the answer came from and how recently it was validated.
When Asterion doesn’t know
Asterion says so plainly. He does not guess, does not claim certainty beyond what evidence supports, and does not fill silence with invention. The response names what’s missing from Canon so you know exactly what the gap is.
When Asterion knows something adjacent
Asterion surfaces what Canon does contain that’s related, explains why it’s related, and points toward where the specific answer you want can actually be found — whether that’s a generation run, a validation report, or a domain that hasn’t been surveyed yet.
How the architecture enforces this
Hallucination prevention isn't a prompt instruction — it's a retrieval architecture. Asterion only answers from items that pass authority, freshness, and confidence gates. Items below threshold aren't surfaced, and the Q&A service is designed to detect empty retrieval results and respond with an explicit acknowledgment rather than a generated approximation.
- →Retrieval-grounded: answers come from Canon, not LLM weights
- →Authority gates: items below approved status are excluded from Q&A retrieval
- →Freshness gates: expired items are excluded or flagged before surfacing
- →Confidence propagation: source evidence confidence chains through to the answer
- →Empty retrieval → explicit unknown: no answer generated when Canon has nothing
The Canon
The Canon is a governed, append-only knowledge store — the source of truth that all three AI identities and human reviewers read from and write to. Items enter as drafts and become active only after validation. Nothing is silently overwritten.
What Canon Stores
- →Architectural decisions and their rationale
- →Physics concepts and simulation methodology
- →Process documentation and runbooks
- →Simulation experiment results with provenance chains
- →Project history and decision audit trail
Knowledge Kinds
- →decision — architectural or product choices made
- →concept — explanations of how something works
- →process — workflows and runbooks
- →architecture — system design documentation
- →requirement — constraints and must-haves
- →experiment — simulation result records
- →configuration — environment and tuning values
Knowledge Lifecycle
Items promoted from experiments enter as candidate_promoted_from_experiment and require human approval before becoming active. Supersession chains are tracked explicitly — no silent overwrites. All transitions are recorded in an immutable audit log.
Retrieval Architecture
Asterion answers questions by retrieving from Canon — not by generating from training data. The retrieval strategy prioritises governance over similarity: authority and freshness filter results before semantic ranking is applied.
Metadata-First Hybrid Search
Every query passes through five ranked signals before results are returned. An item that is semantically similar but stale or low-authority will rank below a fresher, more authoritative item — preventing outdated Canon entries from surfacing ahead of current ones.
Semantic similarity (1024-dimension vector search) is applied only to the pre-filtered, authority-ranked candidate set — not to the full Canon.
The Auditor Corps
Six bounded specialist auditors maintain Canon health. Each has a single clear responsibility, runs on its own schedule, and publishes findings to the evidence ledger — never directly modifying active Canon. Human review is required before any auditor finding reaches active status.
Survey
Every 60 minDiscovers new knowledge from configured data sources and creates draft Canon entries for human review.
- →Scans universe database (galaxies, star systems, planets)
- →Scans codebase modules, API routes, and configuration
- →Deduplicates against existing Canon before creating new items
Validity
Every 30 minCompares canonical truth to observed reality — detects stale entries and supersession conflicts.
- →Flags items not validated within 30 days
- →Detects supersession conflicts (active items with broken supersession links)
Pattern
Daily 2 AMDetects Canon structure anomalies and galaxy physics outliers. Auto-promotes findings with Measured (deterministic) confidence.
- →Canon structure: orphaned items, unlinked clusters, repeated patterns
- →Galaxy physics: virial ratio, spectral-type distribution, stellar mass outliers
- →Promotes Measured findings to proposed Canon status; Estimated findings always go to human review
Simulation
On eventValidates client GPU experiment results and decides whether findings are eligible for Canon promotion.
- →Evaluates experiment results: CONFIRMED / REJECTED / INCONCLUSIVE
- →Promotes confirmed findings with Measured confidence to proposed Canon status; Estimated confidence always goes to human review
- →Raises anomaly alerts when validated Canon items are contradicted by sensor data
Archive
DailyFlags stale, superseded, and orphaned items for human review. Never auto-archives — all findings go to the review queue.
- →Items with no validation for > 90 days
- →Supersession chain breaks (A → B → C where B is missing or inactive)
- →Dead galaxy references (items pointing to non-existent galaxy IDs)
Codebase
Nightly (dev only)Validates that code and documentation stay aligned — detects drift between architectural intent and implementation.
- →Module and API docstring accuracy (large language model (LLM) evaluation)
- →Broken file path references in documentation
- →Architecture drift: planning docs vs. deployed CDK stacks
- →Stale TODO/FIXME markers (> 30 days old)
- →Single responsibility violations and cyclomatic complexity > 10
Simulation Provenance
Without provenance, a simulation result is just a number. You cannot reproduce it independently, submit it for peer review, audit it for a grant claim, or trust it for a mission-critical decision without knowing exactly what was running when it was produced. Provenance chains solve this: every output is linked to the precise inputs, physics model version, and generation parameters that produced it — making any result independently reproducible from the chain alone.
Without provenance
Simulation outputs cannot be reproduced by an independent party. There is no basis for peer review, no audit trail for grant reporting, and no way to establish confidence in a trajectory or burn decision beyond taking the platform’s word for it.
With a provenance chain
Any result is independently reproducible from the chain alone — no internal state access required. Peer reviewers, grant auditors, and mission planners all have the same verifiable reconstruction path. This is what distinguishes a scientific simulation platform from a calculator.
Connection to Canon promotion
The Simulation Auditor can only confidently promote an experiment result to proposed Canon status when the full provenance chain is intact. A broken or incomplete chain blocks promotion — keeping unverifiable data out of the knowledge base.
Event-Driven Provenance Tracking
Outputs of many-body gravity (N-body)A calculation that follows how many objects, such as stars in a cluster or moons around a planet, pull on one another through gravity over time.More simulations are verified through event-driven deviation checking. Every simulation result carries a provenance chain linking it to the source inputs, physics model version, generation parameters, and the deterministic random seed that produced it. Confirmed results are promoted to proposed Canon status; results that contradict active Canon items raise anomaly alerts for human review.
- TraceabilityEvery output linked to source inputs, model version, and generation parameters
- Deterministic seedThe physics pipeline uses hierarchical deterministic seeding — same seed + same parameters reproduces an identical universe. The seed is a required part of a complete provenance chain.
- VerificationProvenance chains support independent reproduction without internal state access — only the chain is needed
- Deviation detectionEvent-driven N-body output comparison against expected trajectories
- Grant & peer reviewMeets NSF/NASA reproducibility requirements — the chain is the audit artifact
- Mission safetyIf a simulation informed a trajectory or burn decision, the chain proves what was running when
What Confirmed Runs Enable
The confirmed run corpus isn't an archive — it's an active input to five downstream processes. A provenance chain is what makes a run eligible to participate in all of them.
Natural Language Interface
Asterion answers natural language questions by consulting Canon, then optionally calling tools to retrieve live simulation state. Questions about objects that exist in your workspace are answered from measured data — not inferred from training.
Universe Tools
- →Create generation jobs for new galaxies or star systems
- →Predict orbital events (conjunctions, oppositions, transit windows)
- →Design gravitational capture scenarios
- →Add interstellar visitors to existing systems
- →Fetch validation reports for generated objects
Experiment Tools
- →Save and load simulation experiment records
- →List experiments by galaxy or user
- →Generate system specifications from experiment results
- →Submit GPU client results for Simulation Auditor review
Workspace-Aware Context
Asterion doesn't just answer questions about abstract physics and architecture. He knows what you're working on right now — the scene you have loaded, the galaxies you've generated, the simulations currently running. Every answer is scoped to your universe, not a generic knowledge base.
Example — live workspace response
“What is before us now: Your scene contains a galaxy called ‘New Stellar Cluster’ — currently empty of star systems. The Auditors have brought me detailed measurements of stellar populations across multiple spectral types: from the rare, massive O-type stars (averaging 20.7 solar masses) down through the abundant M-type dwarfs (averaging 0.2 solar masses). These are verified records from the cosmological production authority.”
Asterion knows your workspace
The current scene, loaded galaxies, running simulations, and generated objects are all part of the context Asterion answers from. Asking “what stars are in this system?” returns measured data from your generation run — not textbook averages.
Answers are scoped to your universe
Canon is populated by Auditors surveying your actual generated universe. When you ask about stellar populations, virial ratios, or object counts, Asterion is reporting on measurements taken from your data — not returning generic astrophysics definitions.
This is the distinction between a knowledge platform and a chatbot
A chatbot answers from training weights. Asterion answers from Canon items that were created by Auditors analyzing your generated universe — grounded in your specific objects, parameters, and simulation results. The difference is measurable: every answer carries the source, freshness class, and confidence of the evidence behind it.
What workspace context includes
- →Active scene — the galaxy or system currently loaded in the studio
- →Generated objects — galaxies, star systems, planets, and celestial bodies from your generation runs
- →Running simulations — N-body experiments and mission trajectories in progress
- →Validated measurements — stellar population surveys, virial ratios, and spectral distributions from the Auditor Corps
- →Generation parameters — the seeds, physics model versions, and configuration values that produced your universe
Edge & Air-Gap Deployment
Cosmos Genesis is designed to operate in environments where cloud connectivity is unavailable, unreliable, or prohibited. The Canon export architecture was built with this in mind from the start — targeting space operations, emergency response, and secure facility deployments where sneakernet is the only available data channel.
Signed Canon Bundles
- →Canon exports are Ed25519-signed bundles — integrity is verifiable without a network connection
- →Point-in-time named snapshots enable version-pinned offline deployments
- →Organization-scoped exports carry full provenance chains for every item
- →Bundle verification requires no cloud call — signature check is fully local
Target Environments
- →Deep space missions — intermittent or absent uplink windows
- →Emergency response — field deployments with no reliable infrastructure
- →Secure facilities — air-gapped networks where cloud egress is prohibited
- →Remote operations — areas where connectivity is available only in windows
On-Device / Cloud Latency Indicator
Every Asterion response shows the time spent on-device versus in the cloud — for example, 0.0s on-device · 5.0s cloud. Canon retrieval, validation, and signing are fully local and complete in milliseconds. Cloud latency reflects Bedrock inference time. As local LLM support expands, the cloud figure drops — the indicator makes that progress concrete for every response.
- →Canon retrieval, signing, and validation: fully local — typically 0.0s
- →Embedding generation: cloud (Bedrock Titan) today, local model planned
- →Q&A inference: cloud (Bedrock Claude) today, local LLM runtime in development
- →Indicator is per-response — operators see the split on every interaction
Active development direction. Canon export, signing, retrieval, and validation are production-ready today. Local embedding inference and offline Q&A are in active development — the cloud/local indicator tracks progress toward a fully air-gapped deployment.
Personas
Asterion is designed to surface Canon through five personas — each adapted to a distinct user context. The underlying knowledge base is shared; the delivery would adapt to who is asking.
Mission Controller
Operational context. Focused on mission parameters, orbital mechanics, and real-time simulation state.
Lab Assistant
Technical depth. Answers physics and engineering questions with precision — appropriate for researchers and engineers.
Educator
Accessible explanations. Prioritizes clarity and analogy over technical detail — appropriate for students and general audiences.
Guide
Discovery orientation. Helps users explore the generated universe — surfaces interesting objects, phenomena, and context.
Game Master
Narrative framing. Translates physics data into story-relevant context for game developers and interactive experience designers.
Planned, not yet implemented. The five personas above describe the intended delivery layer. Today, every query is answered through Asterion's single unified interface — persona-specific framing has not been built yet.
AI Engineering Pipeline
How Asterion integrates into the AI-human hybrid development methodology — the three AI identities, Codebase Auditor review, and human architectural oversight.