THE UNIVERSAL DATA SUBSTRATE

AI MEMORY, BUILT AS ONE ENGINE

Semurg builds AI memory: a sovereign, CPU-native universal data substrate that does the work of a graph database, vector index, time-series store, object store, relational store, cache and search index with no third-party storage or query engine anywhere in the data path.

Full read/write: it stores and serves a living knowledge graph on standard CPUs, with no GPU anywhere in the product. On your hardware, in your jurisdiction, down to fully air-gapped.
One engine on the CPU clusters you already run. No GPU fleet, no new estate.

Ask the best AI model in the world about your operation and it answers fluently — about operations in general. Your assets, your procedures, your incident history, the exemption negotiated with the regulator in 2021: none of it is in the model. AI arrives brilliant and blank.

There is a workaround, and your people are already doing it: hunt down the right file and paste it into a chatbot. One question, one file, one upload into somebody else's cloud at a time. The whole folder is a governance breach. The whole database is impossible. And everything you do feed through a model is metered — token by token, your own knowledge billed back to you — until AI operations become their own line in the P&L.

What it's missing is memory — and yours is already written down. A work order from 2019. The email thread where the real decision got made. Scattered across systems, filed under logic that leaves when people do.

Semurg connects to those systems — without writing to them — and turns what they hold into a single graph: assets, documents, people, relationships. Ask in plain English; the answer comes back cited to source, from models inside your perimeter. The meter stays off.
THE UNIVERSAL DATA SUBSTRATE FOR THE CLUSTERS YOU ALREADY RUN

Enterprise applications can spend their first year wiring into your systems.

Ours are born connected. The first two are live.

When someone retires, their thirty years of institutional knowledge doesn't walk out the door. It stays — queryable, cited, and connected to everything it relates to.

Every file. Every format. Every silo. Connected, queryable, and cited to source.

A regulator requests documentation covering four years of maintenance procedures at a generation site. Three people across two teams spend the next four days assembling the response. They search six SharePoint sites, two network drives, and an email archive. They're not sure they've found everything. The senior engineer who managed that site retired last year, and the filing system went with him.

The Semurg Knowledge Graph connects to all of these sources simultaneously — SharePoint, network drives, email servers, legacy databases, cloud storage — without writing to any of them. It ingests every file format: text documents, PDFs, spreadsheets, engineering drawings, scanned inspection reports, presentations, audio, video. It reads them and extracts every entity, every topic and every relationship into a knowledge graph where documents are connected not just by keywords but by meaning.

Any authorised user can then ask a question in plain English. The answer comes back with citations to the specific documents that support it. Not a list of search results. An actual answer, traceable to source. The Knowledge Graph is built to return no answer rather than an unsupported one, and because every answer arrives with its citations, you never have to take it on faith.
  • Connects to your existing systems — without writing to them
  • Every answer cites the source document — traceable, not generated
  • Respects access controls — users see only what they're permitted to see
  • The knowledge graph grows as new documents enter — connections form automatically

Semurg Knowledge Graph

Application 01

Your organisation defines the policies:

Use any AI. Nothing sensitive leaves the building.

Most regulated organisations can't let employees use external AI models because sensitive data might leave the perimeter. SAIG is built to close that path.

SAIG governs what goes out. When employees interact with any external AI model your organisation provides - any provider, any model - SAIG identifies and masks sensitive information before the query leaves the perimeter. Your people get full AI capability. Your data stays governed.

You define the policies. SAIG ships with comprehensive detection defaults covering personal information, financial data, government identifiers, and more. Your admin adds custom categories specific to your operations - internal project codes or infrastructure identifiers, whatever your compliance requires.

For the user, the experience is indistinguishable from using the AI directly. For the organisation, no sensitive data released to a third party under the policies you set.

Every interaction is logged into the knowledge graph. Over time, this builds a continuous picture of how your organisation uses AI — what people need to know, where expertise gaps exist, and which teams are getting the most value. SAIG doesn't just govern AI usage. It makes that usage visible, measurable and auditable.
  • Which AI providers employees can access
  • Which data categories are protected — and what happens when they're detected (mask, flag, or block)
  • Custom sensitivity rules for your industry
  • Deploys on any cloud, on-premise inside your perimeter, or completely air gapped with no external calls at all

SAIG — semurg ai gateway

Application 02
Substrate-native inference: models run where the memory lives
Everywhere else, your data travels to the model: shipped to a cloud endpoint, or copied out to a GPU cluster with its own serving stack. Semurg inverts it. To the substrate, a model is just more data: its weights live in the same store as your graph, shared experts stored once, streaming from disk through the same engine. Native streaming is built for sparse mixture-of-experts models, deliberately: only the experts a token needs are active, so the working set stays small enough for disk to feed just in time.

A model larger than the machine's memory still answers, because disk is the source of truth and memory is just the pipe. It runs on your own CPUs, beside the knowledge it reasons over; want a dense or different model, attach it as a sidecar. No GPU fleet. No per-token meter. Nothing leaving the building. A GPU still wins raw single-stream latency, and we say so plainly: what changes is where the model runs, what it costs, and who controls it.
No Kubernetes, No memory sprawl, No per-token bill.
Agents are only as useful as what they remember, and today each agent stack brings its own memory: a vector store here, a cache there, another copy to secure, another pipeline to sync. Then comes production. One agent is a script on a laptop; a swarm is normally a container per agent, with a Kubernetes cluster to herd the fleet. Semurg needs neither. Agents run as lightweight concurrent workers inside the engine itself, reading and writing against the substrate.

One store means every agent shares the same memory: searchable by meaning, walkable by relationship, with identical facts stored once. One runtime means agents inherit the house rules: the same access controls as everyone else, every action in the same audit trail, and forgetting handled as a retention policy you set, not code you write. With models running locally, an agent's thinking carries no per-token bill: a swarm costs hardware, not API credits. How capable your agents are is your choice of model and framework; the memory, the runtime, the governance and the economics underneath them are the substrate's job.

the model comes to memory.

one memory for all your agents.

Your infrastructure. Your jurisdiction. Your boundary.

Semurg is sovereign AI infrastructure in the literal sense: the memory runs where the data already lives.
The sovereignty is architectural, not contractual — one engine with no third-party storage or query engine in the data path,
no GPU fleet forcing you into someone else's cloud, nothing that structurally requires data to leave. Useful for any organisation.
Decisive for the ones that can't send their data anywhere: energy, government, financial services, mining, critical infrastructure.

On-Premise.

Runs on your hardware, in your data centre. Standard server infrastructure. The substrate, the AI models, and the knowledge graph all operate inside your rack. Nothing leaves. Nothing depends on external connectivity.

Sovereign Cloud.

Deploys within Australian or regional sovereign cloud environments under your chosen jurisdiction. Your data sits under the jurisdiction you chose - not one you didn't.

Air-Gapped.

For Defence, intelligence, and classified environments. Fully isolated deployment with zero external connectivity. Every capability — ingestion, extraction, querying — runs locally as a self-contained system.

SIX questions. answered plainly.

QUESTION 1:

Is Semurg a database?
Yes. One engine - the substrate - does the work of ten kinds of database: graph, SQL, key-value, document, search, vector, OLAP,
time-series, streaming, and object storage. Full read/write: it stores, indexes and serves the knowledge graph as the system of record. Connections to your existing systems are read-only by default: Semurg writes to its own substrate, not to yours. When a workflow needs to write back, you open that path deliberately, under the policies you set, and every write is logged.

QUESTION 2:

What do you mean by substrate?
The layer everything else runs on. AI normally needs three separate stacks: a database estate for the data, a serving stack for model inference, another layer for agents, each with its own storage and its own glue. Semurg collapses them into one engine: the ten database workloads, streaming model inference, and an agent runtime, all CPU-based, on one store. One store is the point: every layer reads the same memory, inherits the same permissions, writes to the same audit trail.

QUESTION 3:

Does our data leave our environment?
No. The substrate, the graph and the models all run inside your perimeter. The only exception is one you create on purpose: route a query to an external model, and SAIG tokenises sensitive information before it leaves, then restores it when the answer returns.

QUESTION 4:

Whose AI models are these,
and do they learn from our data?
Open-weight models run natively inside the substrate, on your hardware: inference happens in the same engine that holds the graph, not in a separate serving stack. Prefer a specific model? Attach it as a sidecar. And Semurg doesn't train models on your data: your knowledge stays in
the graph, as data you govern, not as weights in a model.

QUESTION 5:

Does it replace our existing systems?
Only if you want it to. Day one, Semurg runs alongside, connected read-only to everything you already have: your ERP keeps managing transactions, your SharePoint keeps storing documents, and the substrate holds the one graph of what they collectively know. Nothing
has to be ripped out, nothing has to be migrated.
But the substrate is a full data engine in its own right, so when a system reaches end of life, its workload can move onto Semurg instead of onto another licence.

Question 6:

Where's the proof?
The engine and its architecture live at semurg.io. Benchmarks, reference hardware and reproduction methodology at data.semurg.io. And you can run it yourself: a free single-node install at one.semurg.io, on hardware you already have.
ARTIFICIAL INTELLIGENCE DEVELOPMENT
ENERGY AND UTILITIES
GOVERNMENT AND DEFENCE
Financial Services
LEGAL
MINING AND RESOURCES
DRONES AND ROBOTICS

verticals

The first verticals in line to benefit from the substrate; far from the only ones.

Where the cost of not knowing is not theoretical

The product's whole backend — vector, graph, cache, model serving, agent memory — as one store on one box. Ship the product, not the plumbing.
Regulatory evidence assembled with citations instead of week-long hunts. Maintenance procedures aligned to licence conditions. AI usage policies enforced at runtime.
Sovereign document intelligence for classified environments. Cross-agency knowledge without data movement. Fully air-gapped where required.
External models without exposing client data. Evidence compiled across every silo, cited to source.
For quant and HFT desks: full tick history on hardware you own, queryable past where memory-bound tools stop.
Precedent and clause intelligence across agreements, opinions and filings. Evidence packs with the chain of citation intact.
The conveyor delay, the rerouted train, the demurrage fee: one traceable chain instead of three systems and a surprise. Parts provenance across every rebuild.
Fleet telemetry, mission history and world state in one live graph, on hardware that travels with the operation. Built for the autonomous era, fully air-gapped in the field.

WHERE THE SUBSTRATE APPLIES

We’d rather show you than tell you.

Bring a scenario your organisation faces today — a document that should be findable but isn’t, an operational question that requires three phone calls to answer, a compliance request that takes a week to assemble. We’ll connect Semurg to representative data and demonstrate the answer live.

No slides. No scripted demos. Your scenario, your data shape, your questions.