Models
Two models. Zero guesswork.
Echo when you need speed. Horizon when you need depth. Both with 1M tokens of context, both on one transparent rate card, both through a ZDR inference provider. You pick the model; the price is never a mystery.
Echo
The model you reach for a hundred times a day. Low latency, dense output, and the lowest list price in the catalog. Autocomplete-fast for the small things, careful enough for the medium ones.
Horizon
The model for the problem that has been open for three days. Long-horizon reasoning across dozens of files, hard debugging, and migrations you would not trust to anyone junior. Its own list card, not a multiple of Echo.
Rate card
The whole card. Two lines.
No tiers within tiers, no regional multipliers, no fine print. The full 1M window costs the same per token as the first thousand. Code/CLI and API prepaid debit this same published list card.
| Alias | Family | Window | Cache read / 1M | Write 5m / 1M | Write 1h / 1M | Input / 1M | Output / 1M |
|---|---|---|---|---|---|---|---|
| echo | Echo | 1M | $0.210 | $0.300 | $0.375 | $0.250 | $0.920 |
| horizon | Horizon | 1M | $0.360 | $0.670 | $0.838 | $0.570 | $1.500 |
Data handling
Your code stays yours.
The inference provider is zero data retention. Faelith keeps its own encrypted, access-controlled record of complete model I/O and the complete chat transcript.
No provider-side retention.
The inference provider processes the request without retaining it. There is one rate card for every key; data handling never changes the price.
Complete I/O retained for 90 days.
Faelith stores the exact model input and output plus the complete transcript in its own S3. This includes system and tool messages, tool calls, file reads, writes, grep results, retries, reasoning output, and assistant responses.