ARVICA / DOCUMENTATION
Inference and image generation
Size the service around latency, throughput and memory.
Describe the traffic
Provide model size, context length, expected concurrency and latency target. For image generation, include resolution, pipeline components and batch size. Memory needs depend on the full runtime, not just weight files.
Validate the service
Use representative requests to check startup time, throughput and output quality. Define authentication, request limits and handling for failed requests before accepting customer traffic.
Production expectations
Discuss uptime, maintenance, data handling and capacity growth in the service terms. Selecting an inference template does not automatically create a managed API endpoint or an availability guarantee.