An organizational assistant often starts with a promising demonstration: a model runs on one GPU and answers the team’s document questions. The problem changes when more people arrive, conversations grow and maintenance can no longer take the service offline. A more expensive card is tempting, but incorrect answers, long queues, exhausted memory and outages are different failures.

There are three separate expansion decisions: change the model, add complete serving replicas or change the hardware supporting each replica. A suitable small model replicated across ordinary hosts can provide more independent capacity and a maintenance path. Other workloads genuinely need a stronger model, more memory or faster communication between GPUs. The LLM guide helps separate these decisions before procurement.

Count work, not employees

A thousand employees do not imply a particular GPU count. Many may ask only a few short questions; one automated process may keep issuing model calls. A person reading a response and a request actively generating tokens also consume different resources. Measure request arrival, input/output lengths and calls per completed task.

Define a service-level objective, or SLO, that can be tested. An illustrative target is “95% of interactive requests start responding within two seconds under the specified load.” The number is not a universal recommendation; it turns “fast” into an observable condition. A nightly extraction job may instead need a given number of correct documents completed by morning.

ObservationDecision it informs
Normal and peak arrivals, including burstsReplica count and temporary capacity
Input/output length distributions, history and documentsMemory and how long each request occupies the service
First-token, total-response and inter-token latency at p95/p99Experience of slower requests
Queue depth and age, errors and cancellationsAdmission policy and capacity shortfall
Correctly completed tasksModel capability versus infrastructure failure
Required service during a host outageSpare capacity and placement

vLLM’s metrics include queue, latency, token and KV observations in supported versions. Combine engine metrics with whole-application timing, since the user waits through retrieval, authorization and other stages too.

Locate the waiting time

In a document assistant, authentication, search, reranking and prompt construction precede generation. A slow database or excessive retrieved context may be the bottleneck. Record stages separately before spending on the GPU stage.

Prefill processes the input; decode produces successive output tokens. A long prompt and a long answer with the same total token count need not behave alike. Removing irrelevant passages and budgeting output for the task can release capacity without changing the model, provided the necessary evidence and answer remain intact.

Continuous batching schedules work during generation rather than requiring every request in a batch to finish before new work enters. Orca is a foundational reference. Greater aggregate output still needs to be evaluated alongside interactive latency. The software table identifies compatible execution paths.

Prefix caching in vLLM reuses computed common prefixes. Its main saving is prefill, not a general acceleration of new-token decoding. It also differs from storing and reusing a finished answer. Either cache must respect the application’s data and access boundaries.

Change the model when capability is the constraint

If answers are correct in quiet periods but slow during peaks, enlarging the model may add pressure. Examine scheduling, distribution and replicas first. Change the model when quality, input modality, required context or reasoning no longer fits the current choice.

Bounded extraction or classification can begin with a specialist or a 2–4B candidate; a direct document assistant can compare an approximately 8B model with smaller and stronger alternatives. These are shortlist sizes, not quality guarantees. The model-size guide and RAG stack article explain how to diagnose the task.

A smaller resident model can simplify replication and leave room for active contexts. Quantizing its weights does not reduce every memory component proportionately: KV remains a separate budget. Successful loading is therefore different from stable serving at the target load.

Routing can reserve a stronger model for identified difficult cases. Include the entire path’s success, latency and cost: trying a small model and then retrying on a large one pays for both. Confident phrasing is not a validated classifier of request difficulty.

Replication and model partitioning are different

Here a replica means a complete serving instance, not a new weight release. Four GPUs each running a complete model can offer four serving units. Four GPUs needed jointly for one partitioned model may offer only one such unit.

vLLM’s parallelism guide separates multi-GPU execution from scaling instances. Tensor or pipeline parallelism may solve model fit; data-parallel arrangements distribute work. Actual process dependencies determine whether two apparent units are independent.

Architecture changePotential benefitWhat does not follow automatically
Complete replicas on separate GPUsDistribute independent requestsEach single request becomes proportionately faster
One model split across GPUsMemory and compute for that modelMultiple replacement instances
Several replicas in one hostMore capacity and some process/card fault toleranceSurvival of host failure
Replicas on independent hostsMaintenance or host-outage capacityIndependence of shared power, storage or networking
Temporary cloud replicasFollow variable demandInstant, guaranteed availability of new capacity

Independent requests need not exchange layer outputs continuously across hosts, although document access and routing still use a network. A partitioned model can be sensitive to interconnect latency and bandwidth. This is why inference replication and a tightly coupled training cluster have different infrastructure economics.

A worked capacity example

Assume one replica handles 40 requests/minute within the required quality and latency on the team’s specified workload. This is an educational input, not a measured claim about RTX 4090 or any model. Peak demand is 90 requests/minute, each replica has its own independent host and shared components are not limiting. The 40-request figure is a usable operating rate, not an unstable saturation point.

Ready replicasCapacity under the assumptionAfter losing one replicaResult at 90 requests/minute
28040Insufficient even when healthy
312080Healthy capacity sufficient; outage capacity insufficient
4160120Arithmetic covers loss of one replica

The initial equal-capacity calculation is ceil(target load / capacity per replica). Add one for loss of one independent replica. Load imbalance, smaller batches after distribution, cold caches and retry bursts can make measured total capacity differ from this arithmetic.

Now put the four GPUs two per host on two hosts. Losing one host removes two replicas and leaves 80 requests/minute: the target fails. Under the same assumptions, three two-replica hosts leave four replicas after one host is lost. Plan against the largest capacity group removed by the failure being designed for. A zone-outage objective requires enough capacity outside that zone.

Availability starts with the capacity left over

A second model helps only if routing reaches it and its dependencies work. Authentication, document retrieval, conversation storage, load balancing and networking can be shared failure points. Conversation state needed after failover should not exist solely in one model process’s memory.

Our Targoman deployment used two active hosts with one RTX 4090 each. During model reload on one host, the other continued translation, summarization and chat, but waiting and some cancellations increased. The original Persian account illustrates the difference between remaining reachable and preserving the same service objective. It is not a quantified availability guarantee.

Kubernetes topology spread constraints can implement placement across hosts or zones. PodDisruptionBudget helps limit voluntary disruptions; it does not prevent sudden host failure. Capacity and application behavior must make the placement useful.

For planned replacement, load and warm the new instance before admitting traffic, then drain the old one. Sudden failure can interrupt an in-progress stream. Retrying an agent operation must not duplicate an external payment, document creation or other side effect; operation identities and duplicate handling belong in the application. Request-path availability is not uninterrupted continuation of every generated response.

When a larger GPU addresses the actual problem

More VRAM matters when the required model, context or KV cannot fit. More per-instance capability matters when an isolated request is already too slow; replicas alone do not fix that latency. A professional platform can also reduce rack space, energy per useful job or operating effort for sustained demand.

The GPU guide compares memory, interconnect and platform requirements. Compare the complete architecture at the same quality and latency, including the additional host needed for availability. One powerful card does not create a second independent instance. Two 24 GB cards also do not automatically become one unified 48 GB memory space; the 24/48 GB article explains the distinction.

Billing flexibility is different from autoscaling

Temporary capacity is useful while demand and model choice are uncertain, for experiments or for periodic peaks. Once the pattern is known, baseline demand can use an appropriate purchase, reservation or rental arrangement while variable work uses another route.

Pay as you go describes billing; autoscaling describes capacity control. Closing an application need not stop a rented machine’s bill. EC2’s lifecycle documentation distinguishes running instances from stopped ones and separately billed resources. Other services meter worker startup, execution or an idle timeout. Compare the actual billable unit and when it starts and stops.

GPU sharing is another allocation choice: services can receive managed portions rather than each reserving a whole card. For modest or variable demand, this may avoid paying for unused exclusive capacity. Memory allocation, interference, isolation and response time under simultaneous demand remain part of the comparison. Sharing is not synonymous with an interruptible instance, nor does it necessarily require the same virtualization mechanism on every platform.

Targoman states that it is the only provider in Iran offering the combination of GPU Sharing and pay-as-you-go billing described here.

Capacity arrangementUseful forCost or operational condition to include
Always-ready GPU instanceBaseline interactive demandIdle time, spare capacity and management
Managed shared GPU resourcesLower or variable demandMemory share, concurrent-use rules and busy-period latency
Temporary replicas beside a baselinePredictable or seasonal peaksProvisioning time, GPU availability and data transfer
Scale-to-zero serviceInfrequent work that tolerates waitingCold start and resources still billed after shutdown
Interruptible/Spot capacityRestartable or resumable queued jobsInterruption and recovery work
Hosted model APILimited demand or a specialized routeQuality, latency, limits and all calls per task

AWS Spot interruptions are one documented example of reclaimable capacity. Such a service needs recovery or an alternative path if the workload cannot simply wait. Shared allocation alone does not imply reclaimability; the service contract determines that property.

A worked example of following the peak

Return to the assumed 40 requests/minute per replica. Suppose normal load is 30, so two ready independent replicas preserve sufficient capacity after losing one. The 90-request peak needs four replicas for the same failure objective.

In an illustrative 720-hour month, keeping four replicas ready continuously uses 2,880 replica-hours. Keeping two ready continuously and adding two for 60 hours each, including preparation and idle shutdown time, uses:

2 × 720 + 2 × 60 = 1,560 replica-hours

That is about 46% fewer replica-hours. In this example each replica has one GPU, so GPU-hours happen to match. A shared GPU or multi-GPU replica changes the unit. It is not automatically a 46% reduction in the whole bill: tariffs, storage, networking and management can differ between capacity types.

Extra replicas must be ready before the peak and placed to preserve the assumed failure tolerance. Starting them only after a queue forms adds provisioning, weight download, loading and warm-up to the wait. Ray Serve’s autoscaling guidance discusses cold starts; infrastructure scaling is another layer. Removing model processes does not necessarily release cloud machines or their billed resources.

For predictable peaks, prewarm capacity. For sudden bursts, retain appropriate ready headroom and admission limits. Queue age and latency approaching the objective, together with memory and in-flight work, can be more informative than GPU utilization alone. Scale-down cooldowns reduce oscillation; minimum fault-tolerant capacity and cost ceilings still apply.

A hybrid local/cloud fallback needs compatible weights, templates, permitted data access and networking prepared in advance. Autoscaling cannot create unavailable regional capacity or quota, and an instance downloading everything during an outage is not equivalent to a warm spare.

An unlimited queue is not a scaling strategy

Separate short interactive work, long analysis and background jobs where necessary. Set input, output and runtime budgets appropriate to each. When capacity is exhausted, bounded waiting or asynchronous completion can be preferable to an ever-growing queue. A lower-quality model should follow an explicit product policy, not silently replace sensitive analysis.

Retries can amplify overload. Limit attempts and use increasing delays with jitter; propagate cancellation to the engine when supported so abandoned work stops consuming capacity. Google’s overload-handling guidance connects admission, retries and load balancing.

Evaluate bursts, long inputs, replica loss, cold-cache recovery and dependency failures. Record late, cancelled and failed requests alongside successful output. Useful capacity counts tasks meeting quality and timing conditions, not every token generated.

Match the change to the symptom

Observed symptomChange to examineDecision criterion
Wrong answers even without loadRetrieval, prompts, a better-suited model or specialist routeMore accepted tasks on representative examples
Good answers but queues during peaksScheduling and additional independent replicasMore capacity within quality and latency targets
A single isolated request is too slowInput, application path, model or per-instance hardwareLower latency for that request
Context/KV exhausts memoryInput selection, placement, cache strategy or more VRAMPreserve required evidence without memory instability
Host loss disrupts the serviceReady capacity in another failure domainSufficient surviving capacity and healthy dependencies
Mostly low demand with peaksSmaller baseline plus timely temporary capacityLower actual total cost while meeting the peak
Stable demand with expensive rentalPurchase, capacity agreement or continued rentalComplete operating, spare-capacity and change costs

Use workload evaluation to define accepted outputs. Divide the period’s complete costs—including idle readiness, failures and escalation paths—by accepted work in that period. This makes the comparison between ordinary replicas and higher-end hardware meaningful.

A small initial service does not need all future hardware on day one. Reproducible serving units, observable expansion thresholds and separate failure domains make gradual growth possible. A larger GPU enters the design when memory, latency or complete-system cost gives it a specific job to do.