BYOK: Per-Request Limits & Key Management
Version: 1.19.2, Date: October 19, 2025
BYOK per-request limits are now enforced, so no single generation can consume an unexpectedly large portion of your credits. This release also gives you visibility into and control over how your keys are being used: status indicators, friendly labels, last-used timestamps, and detailed per-request usage records. Every AI creation workflow now enforces credit checks, model allow lists, and token limits with user-friendly guidance when something needs your attention. The rest of this post is grouped by app area.
Highlights
- Per-request credit ceilings: Prevents runaway costs on any single generation.
- Token input/output limits per plan tier: A safety net against unexpectedly long prompts or responses.
- API key status tracking: Valid, invalid, revoked, or unknown, plus last-validated and last-used timestamps.
- Quota enforcement in every creation workflow: With user-friendly error messages and upgrade paths.
Settings
- New: Per-request credit ceilings on each plan tier. This prevents any single AI request from consuming an unexpectedly large portion of your credits. The ceiling scales with your tier: higher tiers get higher per-request allowances.
- New: Token input and output limits per plan tier. This provides a safety net against accidentally sending extremely long prompts or receiving unexpectedly large responses, keeping costs predictable.
- New: API key status tracking. Your stored keys now have a status indicator (valid, invalid, revoked, or unknown) plus timestamps for when the key was last validated and last used. You can also assign a friendly label to each key (for example “Production key” or “Testing key”) so you can easily tell them apart.
- New: Detailed per-request usage tracking. Every AI request now generates a detailed usage record including the model used, tokens consumed, estimated cost, and variance analysis comparing actual usage against expected. This data powers the analytics that help you optimize your model choices and prompt efficiency.
- New: Comprehensive audit logging. All quota-related events (credit consumption, limit checks, BYOK usage) are recorded in an audit trail, providing a complete history of how your quota has been used.
Across the App
- New: Quota enforcement now runs in every creation workflow. Credit checks, model allowlist enforcement, and token limit validation happen before every AI call across builders, prompt creation, library creation, and version creation. Each check produces user-friendly error messages with clear guidance on how to resolve the issue, including upgrade paths when applicable.
- Improved: Quota-aware model selection. The model selector now factors in your tier’s allowlist when displaying available models, so you only see models you are authorized to use. No more selecting a model and then being told you cannot use it.
- Improved: User-friendly quota error alerts. When you hit a quota limit, a dedicated alert displays your current tier, what limit was reached, and what your options are, including upgrading or adjusting your request. This replaces generic error messages with clear, actionable guidance.

