13+ models. One platform. Zero markup.
Every model on Vultr Inference, with real pricing. No added fees. Bring your own keys.
Built for engineers who need every angle.
3+ models, one prompt, zero waiting.
Fire a single prompt to multiple models simultaneously. Responses stream back in parallel. See where models agree, where they diverge, and catch hallucinations before they reach production.
AES-256 key encryption.
Your API keys are encrypted with AES-256-GCM via the Web Crypto API. Decrypted only in memory during inference. We never log, store, or transmit keys in plaintext.
Ship to production.
Export any response as TypeScript, Python, or cURL — pre-wired with auth headers.
Catch hallucinations.
Cross-examine responses across 3+ models. Consensus scoring flags divergence so you know where to focus review.
Bring your own keys.
Use your Vultr Inference API key. Zero markup on tokens — you pay exactly what the provider charges.
Real-time spend monitoring.
Token counts, estimated cost, and latency per model — all tracked in real time. See exactly what each query costs before you scale.
Watch three models think in parallel.
One prompt, three models, real-time streaming. See responses arrive simultaneously and spot differences instantly.
- Responses stream with live token-by-token output
- Consensus scoring highlights agreement and divergence
- Export any response to production-ready code
- Bring your own API key to get started
Try the API. Right here.
Real code examples for the Vultr Inference API. Edit, run, copy — no setup required.
Powered by Vultr Inference.
One API key, 13+ models from 8+ providers. Zero markup on tokens. Automatic retries across providers.
How it works.
One prompt, encrypted in your browser, fanned out to every model in parallel.
Trusted by developers who can't afford to be wrong.
Engineers using z0.chat to verify critical answers across multiple models.
Trust by design.
Security and privacy aren't features — they're the foundation.
How z0.chat compares.
Multi-model inference platforms compared. No marketing fluff — just the facts.
| Feature | z0.chat | OpenRouter | LiteLLM | Direct API |
|---|---|---|---|---|
| Parallel multi-model inferenceOne prompt → N models simultaneously | ✓ Built-in | ~ Sequential calls | ~ Via router config | ✗ Manual orchestration |
| Consensus scoringAutomatic divergence detection | ✓ Sentence-level | ✗ Not available | ✗ Not available | ✗ Not available |
| Token markupCost above provider pricing | 0% | 5-10% | 0% | 0% |
| Key managementEncrypted storage of API keys | ✓ AES-256-GCM, client-side | ✗ Keys on their server | ~ Self-hosted only | ✗ DIY |
| No backend / no loggingPrompts never touch our servers | ✓ Client-side only | ✗ Routes through backend | ~ If self-hosted | ✓ Direct to provider |
| Real-time streamingToken-by-token output from all models | ✓ Parallel streams | ✓ Single model | ✓ Single model | ✓ Single model |
| Code exportTypeScript, Python, cURL | ✓ All three | ~ API reference only | ~ Python SDK | ✗ Manual |
| Setup timeFrom zero to first query | < 30 seconds | ~2 minutes | ~30 minutes | ~1 hour+ |
| Models supportedVia Vultr Inference | 13+ models, 8+ providers | 200+ models | Any OpenAI-compatible | Varies by provider |
| Cost visibilityReal-time spend tracking per query | ✓ Per-model, per-query | ✓ Account-level | ~ Via logging | ✗ Manual |
Roadmap.
What's built, what's coming, and what's planned. No vaporware.
Core Platform
The foundation is built and working.
Developer Tools
In active development.
Team & Scale
On the roadmap, not yet started.
Model Status Dashboard
Real-time status of all 13+ models. Updated continuously.
Cost Calculator
Estimate your spend. Compare with OpenAI GPT-4 pricing.
Model Leaderboard
Compare all models side by side. Click any column to sort.
| Model ▲ | Provider ▲ | Context ▲ | Input $/M ▲ | Output $/M ▲ | Speed (tok/s) ▲ | MMLU ▲ |
|---|
Prompt Gallery
Curated prompts to get you started. Click “Try it” to launch with the prompt pre-filled.
Zero markup. Pay for what you use.
- 100 queries/day
- 2 models simultaneously
- BYOK with encryption
- Basic code export
- Usage tracking
- Consensus scoring
- Team features
- 1,000 queries/day
- 5 models simultaneously
- Advanced export (TS, Python, cURL)
- Consensus scoring
- Prompt history
- Usage tracking
- Team features
- Unlimited queries
- 5 models simultaneously
- Team workspaces (coming soon)
- Shared prompt library (coming soon)
- SSO / SAML (planned)
- Audit logs (planned)
Questions.
How does parallel inference work?+
Is BYOK really secure?+
What models are supported?+
What is consensus scoring?+
Do you log my prompts or responses?+
What is z0.chat built on?+
Is z0.chat open source?+
Stop guessing. Start verifying.
One prompt to every model. Consensus in seconds. Zero markup, zero telemetry, zero doubt. Free to start — bring your own key.
Launch App →