Under 300MB per model

How TinyAI works.

TinyAI is our family of purpose-built AI models. Each one is under 300MB, runs on standard CPU hardware with no GPU, and deploys entirely inside your environment. Specialised models work together in a three-layer stack to turn data into auditable decisions.

<300MB

Per model

No GPU required

<1hr

Training time

On around 200 labelled examples

<300ms

Inference latency

End-to-end, no round trips

95%+

Accuracy

After fine-tuning on your data

The architecture

Three layers. One decision.

NLP feeds machine learning. Machine learning feeds logic. The orchestration layer coordinates everything, running specialised models in parallel with sub-300ms latency.

L1

NLP foundation

Extraction, recognition and classification. Converts unstructured documents, contracts, forms and scanned records into structured fields the layers above can process.

L2

Machine learning

Pattern recognition, scoring and anomaly detection. Trained on your historical data, around 200 labelled examples in under an hour, with a confidence score on every prediction.

L3

Logic and orchestration

Deterministic decision chains where one model calls the next. Policy rules, routing, thresholds and human-in-the-loop checkpoints, producing one structured decision with a full audit trail.

Deterministic, not probabilistic

Fifteen steps later, still 100%.

Lending decisions are chains. Chain fifteen calls that are each 90% accurate and the end-to-end result falls to about 20%. TinyAI gives the same answer to the same input, every time, so accuracy does not compound away.

  • Same input, same answer, every run
  • Per-token attribution on every decision
  • Model version, inputs and thresholds stored with each record
Explainability

Every decision, explained in two layers.

TinyAI produces structured, deterministic outputs, not free-form text. Every decision is explainable and audit-ready by design.

Layer 1

Structured decision output

A single structured output per model. Each decision carries its confidence score and the features that drove it: input features, weight contributions, model version, timestamp and confidence threshold.

Layer 2

Context attribution

NLP models use contextual attribution. Vision models produce spatial heatmaps. Both operate on the actual model weights, not surrogate approximations, and are stored alongside the inference logs.

Deployment

Three topologies. One performance promise.

All with full data residency and sub-300ms latency. Zero data egress. No PII is ever sent to Synapze or any third party.

01

On-premises

Customer-managed servers. Zero outbound calls. Air-gapped compatible, with full business continuity even when disconnected from the internet.

02

Private cloud (VPC)

Inside your own AWS, Azure or GCP account. No data leaves the cloud boundary. Same performance guarantees as on-premises.

03

Sovereign cloud regions

EU-only or jurisdiction-specific data centres. Supports GDPR Art. 44 to 49 and works with your existing DLP and network controls.

Get started

Start with a one-model pilot.

A 30-minute discovery call, a no-obligation ROI assessment, then a pilot on a single process. Prove it works before you commit.

No GPU procurement · No core banking disruption · Live in 12 weeks