The wire
your AI stack
runs on.

One endpoint. Any model. Full observability. Drop-in compatible with OpenAI — no configuration creep, just results.

<400msp50 latency3+AI providers100%OpenAI compatibleFreeforever1 endpointany modelrequestsZeroconfig driftFullobservability<400msp50 latency3+AI providers100%OpenAI compatibleFreeforever1 endpointany modelrequestsZeroconfig driftFullobservability
How it works

Simple as a wire.
Powerful as a platform.

01

Wire once.

Point your existing OpenAI calls at filament.works/api/route. Add your API key. That's it — no SDK swaps, no schema migrations, no rewriting prompts.

02

Route freely.

Set model to "auto" and filament picks the fastest available. Pin to "gpt-4o" or "claude-3-5-sonnet" for precision. Hot-swap without touching application code.

03

Observe everything.

Every request logged. Latency, model used, token counts, error traces. Query from your dashboard or pull via the observe API — full visibility without extra tooling.

Capabilities

Everything your AI stack needs.

routing

Universal Routing

One endpoint dispatches to OpenAI, Anthropic, Google Gemini, and Venice. Switch models in a single field.

observability

Full Observability

Every call logged with latency, token counts, model chosen, errors. Query via dashboard or REST API.

compatibility

OpenAI Compatible

Drop-in replacement. Zero schema changes. Your existing SDK calls work unchanged.

flexibility

Hot-swap Models

Change your provider without touching application code. Filament resolves "auto" to the fastest available.

tools

Tool Wiring

MCP-ready. skill.md spec. Wire tools across agents and providers through a single protocol layer.

pricing

Free Forever

No usage caps for core routing. Pay providers directly. Filament takes nothing from your inference budget.

Integrate

Three lines.
Any model.

Change your base URL. Add your filament key. Set model to "auto" or any specific provider. Everything else stays the same.

  • Zero schema changes
  • Works with any OpenAI SDK
  • Full observability metadata in every response
  • Instant model hot-swap without code changes
filament.works/api/route
Observability

Every call,
accounted for.

Filament logs latency, model selection, token usage and errors for every request. Pull logs via API or browse the dashboard — no extra instrumentation required.

View your logs
modellatencytokensstatus
gemini-1.5-pro312ms847ok
gpt-4o891ms1,204ok
claude-3-5-sonnet567ms632ok
gemini-1.5-flash189ms412ok
gpt-4o-mini244ms989ok

live request log — last 5 calls

Get started

Your AI stack, finally
under control.

Free API key. OpenAI compatible. Five minutes to production.