Skip to content
Skip to content
● ENGINE LIVESign inCreate an account

wubbery://engine — a CPU-native execution substrate

hundreds of measured modules beneath the operating system. 85 of them are things that don’t happen.

Memory placement, power allocation, capacity planning, inference admission, agent governance — each is somebody’s whole company, and none of them shared a substrate until now. Most of what it does is refuse: a rack that would trip on failover, a compute purchase that returns nothing, a generation that would die at token 400,000. Refusals are hard to fake and easy to demonstrate, which is why we count them instead of averaging them.

The primitive underneath is a decision costing 3.67 µs on one core — no GPU, no network hop, no model call. That is the reason it can sit in front of everything; it is not the product.

OFFICIAL WUBBERY SUBSTRATE ENGINE SPECIFICATIONS & CAPABILITY MATRIX

WUBBERY Engine vs. Legacy Cloud AI Stack

Measured on an 11th-gen i5-1135G7 laptop, single-threaded, best-of-5 over 100,000 decisions. Re-run any row yourself: GET /v1/bench/overhead.

CANONICAL ROUTING SPEED:3.67 MICROSECONDS272,307 Decisions / Sec
ROUTING DECISION COST:NO MODEL CALL32.9% Measured Cost Reduction
SHARED MEMORY REQUIRED:20.3× SMALLERMeasured across 64 hosts
Engine Capability⚡ WUBBERY Substrate EngineLegacy Cloud AI SaaS StackEnterprise Business Impact
Canonical Decision Latency
Category: throughput
3.67 μs (0.00367 ms)
1,800 ms - 2,400 ms
27,248x FASTER THAN A 100ms MODEL CALL IN THE PATH
Single-Core Decision Volume
Category: throughput
272,307 Decisions / sec
0.4 Decisions / sec
HIGH-DENSITY RACK THROUGHPUT
Realtime Viseme Mouth Alignment
Category: throughput
< 15 ms (60 FPS Lip-Sync)
750 ms - 2,200 ms Lag
ZERO SPEECH-TO-TEXT BUFFERING
Per-1k Turn Token Royalty
Category: cost
Price delta between tiers
$12.00 - $120.00 / 1k Turns
32.9% MEASURED COST REDUCTION
Enterprise Data Centre Cost
Category: cost
32.9% measured on real replayed traffic
Every request at the top tier
69% VS AN ALWAYS-PREMIUM BASELINE
Facility Energy & Power Draw
Category: esg
Inference energy falls with the routed share
All inference at premium tier
DEPENDS ON YOUR AI SHARE OF IT LOAD
Statutory ESG Filing Compliance
Category: esg
Local audit log of every routing decision
Provider-side logs you do not hold
APPEND-ONLY JSONL, TIER LABELS ONLY
On-Device Memory Encryption
Category: security
AES-GCM 256; keys never leave the process
Cloud Gateway Telemetry Logs
LOOPBACK-ONLY BIND, NO EGRESS
Algorithm Code Protection
Category: security
Server-only core; never shipped to a client
Exposed API Keys & Plaintext
COMPILED, NOT OBFUSCATED — SEE /SECURITY
Multi-Platform Delivery Formats
Category: delivery
Anthropic + OpenAI dialect gateway; Docker
Restricted SaaS HTTP Endpoint
RUNS IN YOUR VPC, BINDS LOOPBACK
Executable Terminal Daemon
Category: delivery
Single Node process, one env var
Complex Cloud Setup
ANTHROPIC_BASE_URL=127.0.0.1:8787
ENTERPRISE DATA CENTRE INFRASTRUCTURE PROFILES

Facility Scenarios & WUBBERY Substrate Impact

Pick a facility size to see what the measured reductions come to at that scale. These are scenarios, not customers — the sizings are illustrative, and the savings are our benchmarked ratios applied to them.

Server Rack Capacity:45,000 SERVERS4,500 Physical Racks
Facility Power Draw:95 MW49.7 MW of compute load freed
Annual Inferences:1.20B / yrHigh-Volume Workload
Net Annual Savings:$149.6M / yr84.54% composed — measured, zero under-served
AVAILABLE WUBBERY ENGINE MODULES & LAYERS

Complete Substrate Modular Architecture

The engine's layers. Routing, memory and prefetch are live and benchmarked; the media layers are roadmap and labelled as such.

@wubbery/engine-video60 FPS (16.6ms frame budget)

Realtime Generative Video & Motion Layer

Generative Video (roadmap)

Roadmap. Not shipping today — there is no released package for it, and the site should not imply otherwise.

Delivery Formats:
Roadmap — not yet released
Offloads 100% of generative video rendering to end-user devices.
@wubbery/engine-viseme< 15ms Latency (60 FPS Visemes)

Sub-15ms Acoustic Viseme Classification Layer

Real-Time Web Audio FFT Engine

Performs real-time frequency domain (FFT) mouth shape classification from audio streams directly inside the browser audio worklet.

Delivery Formats:
NPM PackageAudioWorklet ProcessorC++ Header
Eliminates server-side speech-to-text API calls for lip synchronization.
@wubbery/engine-voice< 45ms Local Audio Synthesis

Client-Side Voice Imprint & TTS Synthesis Layer

Neural Voice (roadmap)

Synthesizes natural neural speech from text using lightweight ONNX models running inside browser WebWorkers.

Delivery Formats:
NPM PackageONNX Model FilePython Package
Reduces cloud text-to-speech API bandwidth by 98%.
@wubbery/engine-router3.67 μs (272,307 Decisions/sec)

Governed Intent Routing & Model Ensemble Layer

Substrate Autonomy Router (9 Lanes)

Routes prompts dynamically across whichever frontier models you already use, based on cost, latency, and permission boundaries.

Delivery Formats:
NPM PackagePython SDKDocker Container
Executes 272,307 routing decisions per second on a single CPU thread.
@wubbery/engine-vault< 4ms AES-GCM Encryption

Sealed Encrypted Memory & Patent Enclave Layer

Harkin Theorem Sealed Security Vault

Enforces strict on-device memory encryption, sealing Harkin Theorem patents, private memory keys, and enterprise customer data.

Delivery Formats:
Roadmap — not yet released
Guarantees zero security leaks or unauthorized data extraction.
wubbery/engine-sidecar< 28ms IPC Gateway

Native Runtime & Audit Ledger Layer

Native C++ / Python Runtime Daemon

High-density native runtime daemon for enterprise data centres and desktop operating systems with cryptographic audit logging.

Delivery Formats:
Docker ImageKubernetes Helm ChartDebian / RPM Package
Integrates directly into Kubernetes clusters and high-density server racks.
YOUR NUMBERS

Put In Your Figures. See What The Measured Ratios Do To Them.

We do not know what you spend, so we do not guess. You supply the baseline; we supply the reductions we have measured, and every line below names the benchmark it came from.

READING ENGINE…
Composed cost reduction84.54%0 under-served
Composed compute saving54.5%0 violations
Guards780 of 780across 190 guards
ComputedPublishedengine unreachable

Typical corporate facility overhead means every watt of compute removed takes roughly half a watt of cooling and distribution with it.

Picking an industry only changes the starting numbers below and the assumed facility overhead. The measured reductions are the same for every industry on this list — they are a property of the engine, not of who is running it.

Your baseline
$/ yr

Annual spend on inference that could be routed. Not your whole cloud bill.

Servers in your own facility running this workload.

W

Your average draw per server. Ours is not a substitute for yours.

$/ kWh

Your contracted electricity price.

×

Facility overhead (PUE): total site draw ÷ compute draw. 1.0 shows compute alone. Yours, not ours.

Inference spend avoided$10.14M / yr84.54% composed cost reduction, measured on the published composed cost benchmark.
Compute energy saved10,742 MWh / yr54.5% composed compute saving, measured on the published composed compute benchmark with zero violations. Compute draw only — not facility power.
Energy cost avoided$1.50M / yrCompute energy saved at your own $0.14/kWh.
TOTAL SITE ENERGY AVOIDED16,113 MWh / yrCompute energy saved × your PUE of 1.50: removing compute removes the cooling and distribution that served it. ESTIMATED — the multiplier is your facility's, the compute reduction is measured.
TOTAL POWER BILL AVOIDED$2.26M / yr16,113 MWh at your $0/kWh. Includes the cooling and distribution avoided, not just the compute.
Total annual saving$11.65M / yrInference spend avoided plus compute energy cost avoided. Both are money not spent — they share a unit and may be added.
Capacity gained, valued$14.37M / yr54.5% compute saved ⇒ 2.20× PER UNIT OF MODELLED COMPUTE, valued at your own inference rate. Not tokens per second and not a hardware measurement: the saving is measured in relative compute units against a declared model, and this is that same measurement read the other way — which is why it is shown separately and added only once.
Combined value$26.02M / yrMoney not spent plus work now doable without buying hardware. Added once, as value — never summed into a single percentage.
WUBBERY share (25%)$6.51M / yr25% of the combined value above — the rate falls from 35% because the performance modules widen what it is charged on. If nothing is saved and no capacity is freed, nothing is owed.
Net to you$19.52M / yrCombined value less the 25% share.
The number we could have shown you instead

$14.00M / yr

Adding the levers instead of composing them on the remainder gives 116.64% — more than the spend itself. We publish the gap rather than the bigger number.

What this is and is not
  • These are estimates from your inputs, not a measurement of your systems.
  • The ratios are measured on our benchmark traffic. Your mix is not our mix — the shadow replay runs against your own logs and returns your number, without changing what you serve.
  • Compute energy is compute draw only. Cooling, PUE and your AI share of IT load are inputs only you have.
  • Cost avoided and capacity gained have different denominators, so they are never combined into a single percentage. They are added once, as money, only to price the share — and both halves stay on screen so the addition can be checked.
  • Capacity is a modelled figure: the compute saving is measured in relative compute units against a declared model, not in tokens per second on hardware.
Want your real number rather than an estimate? The shadow replay runs against your own logs and changes nothing you serve.
Check the ratios yourself: GET /api/wubbery/proof — no key, no input, computed on request.
PICK YOUR MODELS

414 models. Pick any of them to compare with vs. without WUBBERY.

The full routable catalogue across 51 providers, pulled from the OpenRouter public model API on 2026-08-20 — prices and context windows are theirs, not ours, and not retyped. None of these carry a WUBBERY benchmark yet: pick the ones you actually run and we’ll measure those.

Showing 24 of 414
Get API key
Get API key
Get API key
Get API key
Get API key
Get API key
Get API key
Get API key
Get API key
Get API key
Get API key
Get API key
Get API key
Get API key
Get API key
Get API key
Get API key
Get API key
Get API key
Get API key
Get API key
Get API key
Get API key
Get API key
Running something not on this list?Tell us and we’ll add it — free
LARGE DATA CENTRE & ENTERPRISE CAPACITY COMPUTATION

Enterprise Financial & Energy ROI Engine

Enter your exact data centre server rack count, monthly API volume, and energy costs to compute exact mathematical savings.

REF BENCHMARK: ASUS ZENBOOK 13 (SINGLE-THREAD CPU)
Active physical server rack capacity
340.0M turns per month
Current cloud API cost per turn
Data centre utility rate per kWh
EXACT MATHEMATICAL SAVINGS COMPUTATION:NET SAVINGS: $165,361,253 / YR
Annual Dollar Savings:$165.36M / yrReduces annual API & compute spend from $213.46M down to $48.10M.
Compute Energy Saved:182,613 MWh / yrCompute draw only, at the measured 54.5% composed saving. Cooling, PUE and your AI share of IT load are inputs only you have — so this is not a facility power figure.
Latency Reduction:96% FASTERDrops latency from 2,400ms cloud lag down to 85ms on-device speed (ASUS Zenbook 13 ref).
IN-PAGE SUBSTRATE ENGINE ASSISTANT

Ask WUBBERY Engine Assistant

Inquire about Hyperscale AI facility's routable spend, the memory index, or how to run the benchmarks yourself.

Sub-4μs Response Latency
WUBBERY Engine

Hello! I am the WUBBERY Substrate Engine Assistant. How can I help you analyze Hyperscale AI facility (ARCHETYPE — NOT A REAL ORGANISATION) or test our 3.67μs intent routing layer?

MODULAR ENGINE DELIVERY ARCHITECTURE

Deliver WUBBERY Engine Modules Anywhere

Integrate standalone WUBBERY modules into your web apps, mobile builds, or enterprise on-prem containers.

6 Standalone Modules Available

Generative Video Canvas Engine

Renders dynamic lip morphing, breathing sway, and facial deformation on native HTML5 Canvas & WebGL without sending video frames over the wire.

Integration Snippet (NPM)
import { ArteiGenerativeEngine } from "@wubbery/engine-video";

const engine = new ArteiGenerativeEngine(canvasElement);
engine.updateViseme({ mouthOpen: 0.85, mouthWidth: 0.4 }, { fps: 60 });
SUBSTRATE ENGINE ENTERPRISE ACCESS & EMAIL VERIFICATION

Request Substrate Engine Enterprise Access

Create your enterprise profile and verify your business email to unlock sub-4μs intent routing access.

STEP 1 OF 3