DOCUMENTATION

Billing & Payments

Complete guide to billing for Elastic (prepaid) and Dedicated (weekly post-pay) inference.

Billing & Payments

Hypervize uses dual billing on the same account:

Product / chargeSystemFunds
Elastic inference (tokens, embeddings, elastic images)PrepaidBalance
Platform tools + webhook tools Hypervize executesPrepaidBalance (including when the model is on dedicated)
Chronos runs (fee + tools + model steps)PrepaidBalance
Dedicated GPU capacity (node hours, add-ons, HF launch fees)Weekly post-paySaved card

Dedicated capacity never debits prepaid. Managed tool calls always debit prepaid, even on dedicated endpoints.


Elastic Inference (Prepaid Balance)

Elastic inference (chat completions, embeddings, and image generation) is metered by tokens (or images) and charged against your prepaid USD balance.

Embeddings are billed on input/prompt tokens only.

How Elastic Billing Works

  1. You add money to your balance by buying a credit pack through checkout (one-time purchase, minimum $5).
  2. (Optional) Turn on Auto top-up — we will automatically charge your saved card when your balance drops below a threshold you choose.
  3. Free queries (see below): allowed even with a $0 balance. No usage is recorded and nothing is deducted.
  4. After free queries are exhausted (or if you have a paid account), we check your prepaid balance before allowing the request.
  5. Successful requests record usage. We deduct the exact cost (based on our published per-token prices) from your balance.

Token usage for every response is included in the final chunk of the stream so you can track costs in real time on your side.

Model fallbacks and metering

If Hypervize serves an alternate catalog model because the primary provider was rate-limited or unavailable (see Elastic Inference):

  • Usage is recorded against the model that actually ran.
  • You are charged that model’s catalog rates — not the rates of a model that did not serve the request.
  • The response model field should match what was served; use it for client-side cost accounting when present.

Free queries

  • Default: new accounts start with 3 free queries.
  • Chat first-open (optional bump): the first time an unpaid user opens Hypervize Chat, the product may grant +10 free messages (one-time). That is in addition to any remaining free queries from signup — it is not a replacement for the default of 3, and paid accounts do not receive the Chat bump.
  • Free queries cover elastic-style usage for unpaid users; once the pool is empty, prepaid balance is required.

Chronos (scheduled recipes)

Each Chronos run is billed from the same prepaid balance:

TEXT
$0.01 Chronos orchestration fee
+ catalog price of every tool step that executes
+ token pricing for every model compose step

Create and manage jobs with your API key or Chat — see Chronos. Failed mid-run steps may still have charged tools/tokens for work already done.

Adding Funds

Go to Dashboard → Settings → Billing.

  • Click Add Funds.
  • Choose any amount (≥ $5).
  • Complete checkout securely.

The money is added to your prepaid balance as soon as the payment succeeds.

Auto Top-up

On the same billing page you can enable automatic top-ups:

  • Choose a low-balance threshold (e.g. $5)
  • Choose the amount to add (e.g. $10)

When your balance falls below the threshold we attempt a top-up using your saved payment method.

You can see a history of recent auto top-up attempts (both successes and failures).

Viewing Balance and History

Your current prepaid balance is shown prominently on the Billing page.

You can see individual deductions, purchases, and auto top-up activity in the history sections on the same page.


Dedicated Endpoints (Weekly Post-Pay)

Dedicated endpoints run on private GPU capacity that you control. They are billed weekly on your saved payment method — not from prepaid. Product how-to: Dedicated Inference; create flow: Provisioning.

How Dedicated Pricing Works

We calculate charges from actual capacity data.

Compute hours

  • Only time the endpoint was actually running (not SLEEPING / hibernated) is counted.
  • If an endpoint had any activity during the week, we apply a minimum of 1 hour for that endpoint.
  • With scale-to-zero (Min Nodes = 0) the endpoint can sleep when idle. You only pay for the time it was up. The first request wakes it (WAKING UP in the dashboard).

Add-ons (pro-rated by how long the endpoint product existed that week)
Add-ons apply while the endpoint exists in your account (including scale-to-zero when SLEEPING):

  • Logging & telemetry: +$0.15 per hour of presence
  • Custom domain: +$0.10 per hour of presence

Presence is wall-clock time from create until Terminate (or week end), capped at 168 hours per week, with a 1 hour minimum if the product existed at all that week. Terminating mid-week stops further add-on accrual for later hours. A product that exists the entire week (including fully sleeping scale-to-zero) is still charged for the full 168 hours of add-ons.

Terminate
Terminate is available on the endpoint detail page in the dashboard. It removes dedicated capacity and hides the endpoint from your fleet, but does not erase that week’s usage evidence. Compute hours and add-ons already accrued still appear on the next weekly invoice.

HF launch failures
If a Hugging Face model fails to launch we charge a small one-time infrastructure fee:

machine hourly rate × min_nodes (minimum 1) × 1 hour

You see a live cost estimate in the deployment form before you confirm.

Seeing dedicated usage

On Dashboard → Settings → Billing, the Dedicated (weekly post-pay) card shows:

  • An estimate for the current UTC week (compute hours, infra fees, total).
  • The invoice status for last week when a weekly invoice is available.

This estimate does not change your prepaid balance. Final amounts are charged weekly to your card after the week closes.

Payment Method Requirement

To create or run a dedicated endpoint you need a valid default payment method on file. This is indicated on the Billing page as:

✓ Payment method on file (enables dedicated resources)

Weekly Billing & What Happens on Failure

Every week we aggregate usage for any user who had dedicated activity and attempt to charge.

  • On success the week is marked paid and appears on your invoice.
  • On failure:
    • We immediately email you.
    • We mark the week as payment failed and retry for up to 48 hours.
    • After 48 hours with no successful payment:
      • Your dedicated endpoints are disabled.
      • You are blocked platform-wide (API keys stop working for both elastic and dedicated; provisioning is blocked).

When you add or update your default payment method:

  • We automatically try to back-bill the missed week (including any usage during the 48-hour grace period).
  • Once the back-bill succeeds, the platform block is lifted.

Payment Methods

All payments and saved cards are handled through secure checkout and the customer portal.

  • Credit pack purchases (prepaid top-ups) and auto top-ups use checkout.
  • Dedicated weekly charges and back-bills use your saved payment method (off-session).
  • You can view full invoices and manage cards in the customer portal (link available from the Billing page in the dashboard).

"Has valid payment method" only gates dedicated provisioning and running. It does not block your elastic usage.

Platform-wide blocks only occur when you have an unpaid dedicated bill that is more than 48 hours old.


Tool Calls (Platform + Webhook Tools)

Platform tools (e.g. Athena, Vesper, Herald, Pandora, Iris, Ledger) and user-supplied webhook tools that Hypervize runs for you are metered separately from token usage.

  • Every tool execution is recorded as tool usage.
  • These are priced per call according to our published tool rates.
  • Tool call costs are always deducted from your prepaid balance (same pool as elastic).
  • This applies whether tools run during elastic or dedicated inference — dedicated weekly invoices never include tool charges.
  • Chronos step tools and Chronos run fees also use prepaid (see Chronos section above).

Low-balance guard for dedicated endpoints

When using a dedicated endpoint, we require a minimum prepaid balance (currently $0.50) before allowing managed tool execution. If your balance is too low, the tool call is refused and an error is returned to the model as the tool result.

Client-driven (plain) tool calls

When you send plain OpenAI tool schemas (no webhook extension), Hypervize surfaces the tool_calls and you execute them yourself:

  • We do not bill tool execution for those plain tools.
  • You are only charged for the normal inference tokens used in the turn that produced the tool call(s) (prepaid for elastic; dedicated tokens are covered by GPU capacity, not per-token prepaid).
  • Your own execution costs (compute, APIs, etc.) are outside Hypervize.

See Tool Calling and Platform Tools.


Important Rules

  • Prepaid (elastic + tools + Chronos) and dedicated weekly capacity are separate systems.
  • Tools on dedicated still hit prepaid.
  • Dedicated add-ons are pro-rated by how long the endpoint product existed that week (full 168 hours if it existed all week, including SLEEPING scale-to-zero time).
  • The 1-hour minimum for compute only applies to endpoints that actually ran during the week.

Viewing Usage & History

  • Prepaid balance & elastic history: Dashboard → Settings → Billing
  • Dedicated weekly records: Invoices in the customer portal + the dedicated card on the Billing page
  • Per-endpoint burn: Endpoint detail page in the dashboard
  • Tools / Chronos: Usage views and filters under Settings → Usage where available


Questions?

For billing questions reach out to support@hypervize.tech.

Was this helpful?Send feedback