SOVINFRAPrivate modeless inference · EU-hosted.
Sovereign AI stackSEPTEMBER 2026

Inference built around your workload.

Choose a model, let SovAPI route it, or deploy a specialized model for your task. Dedicated European capacity at serverless pricing. Pay per token.

Try SovInfra →
QWEN 3.8-27B FP8 · H200Code, reasoning, vision and tool calling
GEMMA 4 31B BF16 · H200Text, vision and extraction
WHISPER LARGE V3FR / EN audio coverage
BGE-M3Embeddings and semantic search
KOKOROFR / EN speech synthesis

Status

First token from ~50 ms in Europe · Redundant hosting designed for 99.9999% availability

Service status →

Arena

Compare AI. Measured live.

Launch Arena

Built for production AI

One URL. The right model. One meter.

A unified inference layer for direct models, SovAPI routing and specialized workloads — served on European infrastructure operated by SovInfra.
01

Migrate in one line

Point any OpenAI-compatible client to SovInfra. Call a model directly or use SovAPI routing without a custom integration.

02

Operated & specialized models

Gemma 4 31B, Qwen 3.8-27B, Whisper, BGE-M3 and Kokoro are tested, monitored and served by SovInfra. Specialized open-weight models can be deployed for your workload.

03

Test it live →

Run your own prompt on SovInfra and see TTFT and generation speed without creating an account.

04

European by design

No intermediary, no data retention and no training on customer data. Built for repetitive, context-heavy agent loops.

Dedicated. Pay per token.

Priority and high-performance capacity built around your workload. OpenAI-compatible API.

Automatic routing

Let SovAPI choose
"model": "sovapi"
Automatic routing across Gemma 4 31B and Qwen 3.8-27B. Whisper, BGE-M3 and Kokoro are also available through SovAPI.

Direct model call

Gemma 4 31B
"model": "gemma-4-31b"
Text, vision and structured extraction.

Direct model call

Qwen 3.8-27B
"model": "qwen3.8-27b"
Code, reasoning, vision and tool calling.

Benchmarks

Results published by the model developers, according to their evaluation protocols.

Qwen researchGoogle GemmaAnthropic publicationsDeepSeek publications

Run the Arena

Simple usage pricing

Pricing.

Billed by measured usage, with no commitment. Start with 1B free tokens.

Pay-as-you-go

$0.12/ M tokens in
Your first billion tokens are free.No credit card required.
Gemma 4 31B BF16 · H200
$0.12 in · $0.38 out
Qwen 3.8-27B FP8 · H200
$0.12 in · $0.38 out · $0.04 cached
Whisper Large v3
Transcription, 99+ languages$0.00048 / audio minute
Kokoro
Speech synthesis FR/EN$0.62 / M charactersOther languages on request
BGE-M3
Embeddings$0.01 / M input tokens
Usage
One usage counter · One invoice
Support
Email response < 15 min · 08:00–20:00 CET
Data retention
None

One API. Let the router pick, or call a model by name.

Get your free API key

EU hosting · GDPR native · No data training · OpenAI-compatible API

Built for European workloads

Security. Compliance.

01

EU-hosted

All inference runs on European GPUs. No data leaves the EU.

02

GDPR native

No training on user data. End-to-end encryption. DPA available.

03

No retention

Requests are processed in the EU and are not retained or used to train models.

04

Auditable

Our infrastructure stack is open for inspection on Codeberg. Full transparency on security and routing.

1 billion free tokens

No credit card required

Create API key →