OmniAgent
Hybrid Multi-Capability AI Agent. Local-first inference, confidence-aware fallback to Fireworks AI, and self-verification before every answer.
Hybrid Routing
A task router classifies every request and dispatches it to the right capability automatically.
Local vLLM
Generation runs local-first on a vLLM server for low-latency, on-prem inference.
Fireworks AI
Production-grade hosted models serve as the always-available cloud inference tier.
Confidence-aware Fallback
Low-confidence or failed local responses transparently retry on Fireworks — no request is lost.
Provider-aware Verification
Every answer is verified before it returns, by the same provider family that generated it.
Docker Ready
Ships as a container with server and batch modes — one image for the API and offline runs.
How it works
One request, routed, generated locally when possible, verified before it returns.