ARVICA / DOCUMENTATION

Docs / Workloads

Inference and image generation

Size the service around latency, throughput and memory.

Describe the traffic

Provide model size, context length, expected concurrency and latency target. For image generation, include resolution, pipeline components and batch size. Memory needs depend on the full runtime, not just weight files.

Validate the service

Use representative requests to check startup time, throughput and output quality. Define authentication, request limits and handling for failed requests before accepting customer traffic.

Production expectations

Discuss uptime, maintenance, data handling and capacity growth in the service terms. Selecting an inference template does not automatically create a managed API endpoint or an availability guarantee.

Choose an environment