WebInfer connects to 25+ AI providers. Use any combination of cloud APIs and local models.
Plus Chrome's built-in Gemini Nano, and many more.
Leading AI model providers with flagship models
OpenAI
GPT models via OpenAI API
Anthropic
Claude models via Anthropic API
Anthropic on Vertex AI
Claude models via Google Cloud Vertex AI
Google Generative AI
Gemini models via Google AI API
Google Vertex AI
Enterprise Gemini models via Google Cloud Vertex AI
Azure OpenAI
Azure-hosted OpenAI models
Amazon Bedrock
AWS Bedrock models (Claude, Llama, Titan)
xAI Grok
Grok models by xAI
Mistral AI
Mistral Large and Codestral models
Cohere
Command R+ models with RAG support
Optimized for speed and budget-friendly pricing
Groq
Ultra-fast inference with browser search
Fireworks
Fast LLM inference platform
Together.ai
Wide selection of open-source models
Cloudflare Workers AI
Serverless AI inference on Cloudflare edge network
DeepSeek
Reasoning models with context caching
Cerebras
Specialized hardware acceleration
DeepInfra
Cost-effective model hosting
Focused on specific use cases like search, deployment, or custom models
Access multiple providers through unified APIs
Run models on your own hardware with full privacy
llama.cpp
High-performance local LLM inference with llama.cpp server
LM Studio
Local models via LM Studio server
Ollama
Local models via Ollama
Transformers.js
Browser-native Transformers models with WebGPU
Local Models (WebGPU)
Downloaded models running locally with WebGPU
ComfyUI
Local image generation with ComfyUI workflows
AUTOMATIC1111 / Forge
Stable Diffusion WebUI with extensive customization
Fooocus
Simplified image generation with built-in enhancements
InvokeAI
Professional creative workflow for image generation
Chrome AI (Gemini Nano)
On-device AI with Chrome built-in Gemini Nano model
Text-to-speech, speech-to-text, and audio generation
Image generation and visual understanding
All these providers speak the same language. Switch between them freely, combine local and cloud, or run your own infrastructure.
Move from OpenAI to Anthropic to local models. Same code, different provider.
Nodes connect to each other. Your home server can fall back to cloud when needed.
Run your own gateway. Issue tokens to your team or app users. Keep data on your infrastructure or relay other AI inference nodes.
New providers plug in automatically. Your apps gain new capabilities without code changes.
Any AI service or gateway can advertise its inference capabilities by adding a webinfer.json manifest to their server. This enables automatic discovery and configuration.
Learn How