Myriaxon Lighthouse · early access
Know which of your AI you can trust, and why.
As companies add more models, agents and bots, it gets hard to say what is running, what each one can do and whether it is any good. Lighthouse is an AI management system: one place to check, assess, manage and govern every model, harness, agent, bot and swarm in your company, whoever built it.
One rule runs through all of it: nothing is guessed. A parameter that was not measured is left blank, never shown as a zero. Every score says where it came from, and sample data is always labelled.
The demo is the real interface with fictional data. It runs in your browser and is read-only. Nothing you do in it is sent anywhere.
One place to check, assess, manage and govern.
Four jobs, on one platform, for models, agents, bots, harnesses and swarms alike.
A registry of every model, agent, bot, harness and swarm: who owns it, how risky it is, which tools and servers it can reach, what it costs. Discovery lists the models an endpoint serves, so the ones nobody registered stand out. An existing inventory comes in from a spreadsheet, and every row is checked before anything is added.
Point it at any OpenAI-compatible endpoint, the Anthropic API or a bot behind a plain HTTP API. Fifteen short probes score the 14 IQB-14 parameters. Every reply is kept as proof.
Release gates with rollback. An MCP catalog that scans a server's tools before they can be approved, and a check an agent can call before it uses one. Gateway guardrails, a kill switch for each system, and scoped, expiring tokens for agents.
Risk tiers and a policy decision per system. Reviews, waivers and impact assessments. Controls mapped to NIST AI RMF, the EU AI Act, ISO/IEC 42001 and GDPR. An audit trail in which every entry is chained to the one before.
How a score is made, and why you can check it.
Lighthouse measures what a system says when you ask it, and keeps the proof. It scores on IQB-14, the same 14-parameter standard behind the Myriaxon benchmark.
- ProbeFixed prompts about a fictional retailer go to the system's own endpoint. Lighthouse sees only what the system says, and needs nothing installed in it.
- CheckEach reply goes through deterministic checks: text matching, number extraction, schema validation, ordered rubrics. No model grades another model, so the same reply always gets the same verdict.
- ScorePasses become a score per parameter with a 95% interval. The AXS, from 0 to 100, is the weighted mean of the parameters that were measured. An unmeasured parameter is shown as unmeasured, never as zero.
Every number says where it came from.
Evidence modes are shown beside every score and are never blended into one anonymous number. A system with parameters from more than one source is labelled Mixed, and each parameter still shows its own.
- Live measuredLighthouse sent probes to the system's endpoint.
- Metric feedThe owner's own evaluation job submitted the metrics.
- ImportedA scorecard from elsewhere, such as the Myriaxon harness, with its integrity recomputed.
- Sample dataSynthetic. It can never support an approval or satisfy a measured control.
- Not assessedNothing measured yet. Not a zero, and not a pass.
See it.
Real screens from the live demo, with fictional systems. Open the demo and click through any of them.
Works with what you already have.
Built to be checked, with its limits stated.
- No default password. A fresh install generates one and shows it once.
- Secrets are sealed. API keys are never returned by the API, and never sent to a different host than the one they were saved for.
- One guarded way out. Calls to the systems you assess, tool servers, webhooks and connectors pin the address, follow no redirects and refuse cloud-metadata addresses.
- A strict page policy, and an audit log that detects tampering.
- Two-factor and single sign-on. An authenticator app for any account, with one-time recovery codes, or sign-in through your own identity provider over OpenID Connect. You can see where you are signed in and end the other sessions.
- Backups. By default a consistent copy of the database every day, with the newest seven kept. A restore is a command on a stopped server, and it keeps a copy of what it replaces.
- Black box, text only. It judges a system by its replies. What happens inside it is not observed.
- Deterministic checks. An unanticipated phrasing can be marked wrong. The evidence names the check and what it saw.
- A point in time. A score describes one run. Evidence ages, and Lighthouse flags what is stale.
- Illustrative mapping. The framework mapping helps triage. Your legal and compliance team must validate it.
- Connectors and single sign-on are new. Ticketing, chat and e-mail connectors, and sign-in through an identity provider, are tested against local stand-ins, not yet against every vendor's live service.
- One process, one database. It suits a few thousand systems, not a large estate with many teams at once.
Questions.
Is Lighthouse released?
It is in early access. The live demo is the real interface with fictional data. If you would like to run it on your own systems, request early access and we will reply by email.
How is this different from the Myriaxon benchmark?
IQB-14 is the standard. The Myriaxon benchmark measures an agent running inside Myriaxon, checked against what really happened. Lighthouse measures any system through its endpoint with text-only probes, and can import a Myriaxon scorecard or its audit log when you have them, labelling each by where it came from.
Does Lighthouse send our data to Myriaxon?
No. It runs on your own machine or server and has no outside services. It calls out only where you point it: the systems you assess, and any model, webhook or connector you configure.
Can the demo run a real assessment?
In the demo, Run assessment replays a recorded run so you can see the whole flow. In your own Lighthouse it sends the probes to your system, and shows the plan and the number of requests first, before anything is sent.
How many systems can it hold?
We registered 3,000 systems on an ordinary laptop and ran 30 simulated users who never pause: the median page took 115 milliseconds and nothing failed. It is one program with one database, which is why it installs in a minute. Past a few thousand systems, or with many teams using it at once, it is the wrong shape. Run the included load check on your own machine before you rely on those numbers.
Which systems can it assess?
Any model, agent or bot you can reach through an OpenAI-compatible endpoint, the Anthropic API or a plain HTTP JSON API. For systems you cannot reach, you can import a scorecard or feed metrics instead.
Request early access.
Tell us where to reach you and roughly how many AI tools your company runs. We will reply by email.
By sending this you agree that we may reply to you by email. See our privacy notice.