Papers

Systems Co-Design for Latency, Privacy, and Environmental Efficiency in Edge-Cloud LLM Inference

A Coupled-Bottleneck Analysis of Cold-Start Latency, GPU Multiplexing, Environmental Cost, and Edge Telemetry

Joel Anthony

August 202633 min readSelf-published preprintSurvey

Cite this paper
APA
Anthony, J. (2026). Systems Co-Design for Latency, Privacy, and Environmental Efficiency in Edge-Cloud LLM Inference. uncoalesced. https://uncoalesced.com/researchpaper/edge-cloud-llm-inference
BibTeX
@misc{anthony2026edgecloud,
  author = {Anthony, Joel},
  title = {Systems Co-Design for Latency, Privacy, and Environmental Efficiency in Edge-Cloud LLM Inference},
  year = {2026},
  month = aug,
  howpublished = {uncoalesced},
  url = {https://uncoalesced.com/researchpaper/edge-cloud-llm-inference}
}

Summary

summarize

7,313 words, 12 tables, 7 interactive figures → 6 paragraphs

Scripted animation: the summary is written by the author and is the same every time. No model runs in your browser.

Contents
  1. Abstract
  2. 1. Introduction
  3. 1.1 The Edge-Cloud LLM Inference Lifecycle
  4. 1.2 Conversational Data as the New Fuel
  5. 1.3 Scope and Organization
  6. 2. The Cold-Start Latency Taxonomy
  7. 2.1 Weight Transfer as a Physical Limit
  8. 2.2 The KV Cache Restoration Problem
  9. 2.3 Beyond Weight Transfer: Checkpoint- and CUDA-Graph-Level Cold Start Elimination
  10. 3. The Physical and Environmental Cost of Cloud AI
  11. 3.1 Cooling, Facility Overhead, and Water
  12. 3.2 The Idling Tax
  13. 3.3 Embodied Carbon and Carbon-Aware Scheduling
  14. 4. Hardware-Level GPU Multiplexing and Preemption
  15. 4.1 Sub-Millisecond Soft Preemption
  16. 4.2 Cluster- and Orchestration-Level GPU Sharing
  17. 5. Edge Telemetry: The Privacy-Utility Duality
  18. 5.1 Windows Recall: The Surveillance-Adjacent Failure Mode
  19. 5.2 Apple Private Cloud Compute: Architecture as Privacy Guarantee
  20. 5.3 PRISM: Sensitivity-Aware Routing
  21. 5.4 Privacy-Preserving ML Techniques and the Regulatory Boundary
  22. 6. Quantifying the Macro-Value of Micro-Latency
  23. 6.1 Assumptions
  24. 6.2 Time, Cost, Energy, and Water Savings
  25. 6.3 Reading the Numbers
  26. 7. Conclusion
  27. References

Abstract#

Interactive large language model (LLM) applications split their work across a resource-constrained edge device and a hyperscale cloud back end, and the boundary between the two has become the central design problem in modern AI systems engineering. This paper surveys four coupled bottlenecks in that architecture: the multi-stage cold-start penalty that delays a scaled-to-zero inference server by tens of seconds, and the checkpoint- and graph-level techniques that are increasingly closing it; the electrical and freshwater footprint of the accelerators and cooling infrastructure that keep servers warm instead, extended here to the embodied carbon of the hardware itself and the carbon-aware scheduling techniques that shift load toward cleaner grids; the fine-grained GPU multiplexing mechanisms — from driver-level time-slicing, Multi-Process Service, and Multi-Instance GPU to microsecond-scale soft-preemption frameworks such as Hummingbird, LithOS, and Valve, and the cluster-orchestration layer sitting above them — that determine how efficiently a shared accelerator fleet can be packed; and the privacy-utility duality of edge telemetry, illustrated through the contrasting design choices of Microsoft Windows Recall, Apple Private Cloud Compute, and the PRISM routing framework, and grounded in the underlying toolkit of federated learning, differential privacy, and on-device inference along with the regulatory constraints — chiefly GDPR Article 22's treatment of profiling — that bound how any such system can be built. Rather than proposing a single new system to resolve these bottlenecks, we use the measured performance of already-published techniques — GPU-centric KV-cache restoration (Tutti), microsecond-scale soft preemption (Hummingbird, Valve), and CUDA-graph-level cold-start elimination (Foundry, ServerlessLLM) — to construct a fleet-scale accounting of what their combined adoption is worth: a representative 0.212-second per-query reduction in time-to-first-token, applied across a 100-million-user fleet, yields daily savings on the order of 47,000 hours of user time, 83,000 kWh of electricity, and 400,000 liters of water.

1. Introduction#

1.1 The Edge-Cloud LLM Inference Lifecycle#

The integration of large language models into interactive consumer applications such as personal assistants, enterprise copilots, and real-time dialogue systems has produced a data and computation lifecycle that spans from a user-facing edge device to a centralized cloud data center [1, 46]. Smartphones and laptops are the primary interaction boundary, where user input and local context are first captured [46]. Edge hardware has improved substantially, but state-of-the-art LLM inference run entirely on-device remains constrained by limited matrix-multiplication throughput, tight memory ceilings, and thermal envelopes that a phone or laptop chassis cannot exceed [46]. Production systems therefore favor a hybrid edge-cloud paradigm: privacy-sensitive preprocessing, tokenization, and initial embedding generation happen locally, while the heavyweight decoder computation is offloaded to hyperscale, GPU-dense cloud infrastructure [46].

1.2 Conversational Data as the New Fuel#

The substance of the AI competition has shifted with this architecture. Model training and alignment once relied on static, bulk-collected web-scraped corpora; the current generation of systems instead treats dynamic conversational interaction data as the critical input for continuous refinement and Reinforcement Learning from Human Feedback (RLHF). Every query, prompt sequence, and multi-turn dialogue history carries dense semantic signal about user intent, preference, and behavior. Because platforms retain these histories to iteratively train and tune their models, consumer conversation has effectively become the primary commodity of the AI race, and the edge device that originates it must continuously transmit rich, context-dense payloads to the cloud to keep the hosted model useful. That transmission is also precisely what raises the data-privacy, network-overhead, and security-exposure concerns this paper returns to in Section 5.

1.3 Scope and Organization#

This paper is organized around four coupled technical bottlenecks in edge-cloud LLM inference. Section 2 characterizes the cold-start latency taxonomy that governs how quickly a scaled-down inference server can serve its first token, from container- and weight-loading overhead through the checkpoint- and CUDA-graph-level techniques that are increasingly closing that gap. Section 3 quantifies the electrical power, freshwater, and embodied-carbon cost of keeping accelerators warm enough to avoid that penalty, and the carbon-aware scheduling techniques operators use to shift load toward cleaner energy. Section 4 examines the hardware- and cluster-level multiplexing mechanisms — from driver-level time-slicing, MPS, and MIG through microsecond-scale soft preemption up to Kubernetes-level orchestration — that let a data center pack more useful work onto the same fleet. Section 5 contrasts three real-world edge-telemetry architectures to frame the privacy-utility trade-off inherent in any system that tries to anticipate user intent before a prompt is submitted, and grounds that trade-off in the underlying privacy-preserving machine learning toolkit and the regulatory constraints that bound it. Section 6 uses the measured performance of the techniques surveyed in Sections 2 and 4 to construct a fleet-scale accounting of the time, cost, energy, and water value of closing this latency gap across a hypothetical 100-million-user deployment.

2. The Cold-Start Latency Taxonomy#

In serverless and autoscaled deployments, LLM serving suffers a severe performance penalty on first execution, commonly called the cold-start bottleneck [1, 2]. Scaling a service from zero requires a serialized chain of steps to complete before the first token can be produced: node or virtual-machine provisioning, container initialization, image retrieval from a remote registry, execution-runtime setup, model-weight loading, CPU deserialization, host-to-device memory allocation, weight transfer into GPU High Bandwidth Memory (HBM), and GPU warm-up [1, 4].

The relative size of each phase depends heavily on the inference engine. Runtimes built on heavy Just-In-Time (JIT) compilation, such as vLLM, are dominated by software startup rather than container pulling or weight loading [3]. A standard vLLM boot loads weights, compiles operators with torch.compile, captures CUDA graphs across multiple batch shapes, pre-allocates the key-value (KV) cache, and warms the execution pipeline. That sequence alone consumes 176 to 181 seconds, which makes optimizing container pulling or filesystem decompression worth only about 2% of total time-to-first-token (TTFT) [3]. Minimal C++ runtimes such as llama.cpp invert this profile: container startup takes only 2 to 5 seconds, so the bottleneck shifts to image pulling and weight loading, where filesystem-layer optimizations such as dropping gzip compression, using an uncompressed EROFS layer, or enabling Nydus lazy-pulling with prefetch can save up to 36 seconds, a 26% reduction in cold-start TTFT [3].

Table 1. Cold-start latency by inference engine and storage configuration (Llama 3.1 8B, FP16) [3, 4].
Engine / StorageContainer StartupWeight LoadingGraph Capture / WarmupCumulative TTFT
vLLM — overlayfs, uncompressed180.8 s ± 12.2~8.0 s~12.0 s310.9 s ± 8.9
vLLM — overlayfs, gzip176.0 s ± 3.2~8.0 s~12.0 s317.2 s ± 2.2
llama.cpp — overlayfs, uncompressed2.0 s ± 0.044.2 s ± 0.1~5.0 s102.4 s ± 3.1
llama.cpp — overlayfs, gzip5.3 s ± 0.545.1 s ± 1.4~5.0 s138.8 s ± 0.8
llama.cpp — Nydus, prefetch ON1.4 s ± 0.144.0 s ± 0.1~5.0 s77.8 s ± 0.5
llama.cpp — Nydus, prefetch OFF1.8 s ± 0.144.0 s ± 0.1~5.0 s223.4 s ± 2.9
Chart of Table 1Cold-start time to first token, by phase
Container startup Weight loading Graph capture or warm-up Other (not itemised)

The paper does not itemise this remainder. Values marked ~ in the paper (weight loading and warm-up for vLLM, warm-up for llama.cpp) are approximate.

Source: Table 1 of the paper.

View data as a table
Engine and storageContainer (s)Weights (s)Warm-up (s)Other (s)TTFT (s)
vLLM, overlayfs, uncompressed180.8 s8.0 s12.0 s110.1 s310.9 ± 8.9
vLLM, overlayfs, gzip176.0 s8.0 s12.0 s121.2 s317.2 ± 2.2
llama.cpp, overlayfs, uncompressed2.0 s44.2 s~5.0 s51.2 s102.4 ± 3.1
llama.cpp, overlayfs, gzip5.3 s45.1 s~5.0 s83.4 s138.8 ± 0.8
llama.cpp, Nydus, prefetch on1.4 s44.0 s~5.0 s27.4 s77.8 ± 0.5
llama.cpp, Nydus, prefetch off1.8 s44.0 s~5.0 s172.6 s223.4 ± 2.9

2.1 Weight Transfer as a Physical Limit#

Beyond software overhead, moving model weights from storage into GPU VRAM is bounded by parameter volume and interface bandwidth [4, 5]. A 7B-parameter model in FP16 needs 14 GB of VRAM; a 70B model needs 140 GB [4]. Cold-loading a 70B model on a serverless platform can take 15 to 60 seconds [5]. If weights are pulled from remote object storage over a standard network, the transfer is throttled further: a 4 GB 4-bit-quantized 7B model takes about 8 seconds over a 500 MB/s NFS link versus roughly 1 second over a local PCIe Gen4 NVMe SSD at 3.5 GB/s [4]. Mitigations include safetensors streaming, which overlaps disk reads with host-to-device copies to save 20 to 30 seconds on a 70B model [4]; NVIDIA ModelExpress, which uses peer-to-peer RDMA to stream tensors from adjacent warm nodes instead of local disk [2]; and CRIU-based compressed VRAM snapshots, which can cut cold-start latency to 2 to 5 seconds on compatible systems [5].

2.2 The KV Cache Restoration Problem#

Once a model is warm, managing the KV cache for long sequences becomes the dominant data-movement bottleneck [6, 7, 8]. Cache footprint scales with sequence length, batch size, attention-head count, and layer depth; at a 128K-token context, a single user's FP8 KV cache reaches roughly 20 GB [7]. To manage this footprint, serving systems increasingly offload idle cache pages from GPU HBM to CPU DRAM, local NVMe SSDs, or distributed stores such as Redis or Mooncake [6–10]. LMCache, for example, runs an independent multi-process architecture in which an L1 Manager owns a CPU shared-memory pool while a StoreController asynchronously copies KV chunks to an L2 backend [9, 10].

Restoring that cache from SSD, however, introduces its own latency problem [6, 8]. Because modern engines use page-based memory layouts, a logically contiguous KV cache is fragmented into thousands of small, scattered blocks of 16 to 32 tokens each; restoring a 128K-token context for a 64-layer model requires fetching over 256,000 scattered 80 KB objects [8]. Even with GPUDirect Storage (GDS), the CPU still has to coordinate and initiate every I/O, producing GPU stall times of up to 80% and, in some configurations, making cache retrieval slower than simply recomputing the prefill [6, 8]. Tutti addresses this by implementing a GPU-centric KV-cache object store built on GPU io_uring, letting the GPU issue parallel I/O requests directly and coordinating transfers with compute through a slack-aware scheduler [6, 8, 11]. Relative to GDS-enabled, SSD-backed LMCache, Tutti reduces TTFT by 78.3% under strict service-level objective (SLO) constraints, doubles the achievable request rate, and cuts serving cost by 27%, while achieving near-DRAM restoration speed [8, 11].

Table 2. KV-cache staging latency, 64K context, Llama 3 8B [8].
Cache LocationvLLM v0.12.0vLLM v0.17.0Primary Bottleneck
GPU HBM (native)9.4 s1.7 sMemory bandwidth
CPU DRAM offload (L1)30.5 s14.2 sPCIe Gen4 host-to-device copy
Standard NVMe SSD (L2)76.9 s72.8 sFragmented page I/O, CPU staging
NVMe SSD + GPUDirect Storage73.0 s72.3 sCPU-centric control-path interrupts
Tutti — GPU-centric io_uring~11.2 s~1.9 sBounded by SSD physical read limit
Chart of Table 2KV-cache restore time at 64K context, by storage tier
vLLM v0.12.0 vLLM v0.17.0

At 64K context on vLLM v0.17.0, a standard NVMe restore takes 72.8 s; Tutti takes about 1.9 s.

Source: Table 2 of the paper. Tutti values are approximate (~) in the paper.

View data as a table
Cache locationvLLM v0.12.0vLLM v0.17.0Bottleneck
GPU HBM (native)9.4 s1.7 sMemory bandwidth
CPU DRAM offload (L1)30.5 s14.2 sPCIe Gen4 host-to-device copy
Standard NVMe SSD (L2)76.9 s72.8 sFragmented page I/O, CPU staging
NVMe SSD + GPUDirect Storage73.0 s72.3 sCPU-centric control-path interrupts
Tutti: GPU-centric io_uring~11.2 s~1.9 sBounded by SSD physical read limit

2.3 Beyond Weight Transfer: Checkpoint- and CUDA-Graph-Level Cold Start Elimination#

Sections 2.1 and 2.2 treat cold start as a data-movement problem: getting weights into HBM and a KV cache back into position. Two recent systems show that once those transfers are fast, a third, less-discussed bottleneck dominates — reconstructing the execution graph itself. ServerlessLLM addresses the data-movement side directly with a multi-tier checkpoint-loading system that distributes model checkpoints across underutilized GPU memory and local storage tiers, paired with live inference migration and a startup-time-aware scheduler that decides which tier to load from based on current cluster state; relative to conventional serverless inference baselines, it reports a 6- to 8-times reduction in cold-start latency [61].

Foundry targets a different phase of the same problem. Even after weights are resident, modern inference engines spend minutes capturing CUDA graphs and compiling kernels for every batch shape a serving instance might see — a cost that recurs on every autoscaling event, not just the first cold start. Foundry observes that graphs captured at different batch sizes for the same model share an identical topology and differ only in per-node launch parameters, so it captures execution context once, offline, as a reusable template, then specializes that template at serving time via NVIDIA's cuGraphExecUpdate API rather than reconstructing the graph from scratch. For a large mixture-of-experts model (Qwen3-235B-A22B), this cuts initialization from roughly 10 minutes to 3.9 seconds, a 99% reduction; for a mid-sized dense model (Qwen3-14B), the same technique cuts 36-to-48-second initialization to 1.7–1.8 seconds [62].

Table 3. Cold-start elimination systems compared [8, 61, 62].
SystemTarget BottleneckMechanismReported Improvement
ServerlessLLMCheckpoint loading across memory/storage tiersMulti-tier checkpoint loading + live inference migration6–8x cold-start latency reduction
FoundryCUDA graph capture / kernel compilationTemplate-based graph materialization via cuGraphExecUpdateUp to 99% reduction (~10 min → 3.9 s, large MoE model)
Tutti (Section 2.2)KV-cache restoration from SSDGPU-centric object store via GPU io_uring78.3% TTFT reduction vs. GDS-backed LMCache

Together, the three systems suggest that no single stage — weight transfer, KV-cache restoration, or graph capture — is fundamentally unsolvable in isolation; the harder problem, addressed empirically in none of the three, is orchestrating all three optimizations inside one autoscaling event without their assumptions about resident state (which tier holds the weights, whether a compatible template already exists, whether the KV cache belongs to the same session) silently invalidating one another.

3. The Physical and Environmental Cost of Cloud AI#

Generative AI is frictionless from the user's chair, but the cloud infrastructure behind it carries a substantial thermodynamic and environmental overhead. Dense GPU clusters built for LLM serving are high-intensity thermal systems: an NVIDIA A100 SXM4 has a Thermal Design Power (TDP) of 400 W; the Hopper-generation H100 SXM5 and H200 SXM5 sustain 700 W under active inference; and the Blackwell B200 SXM6 draws 1,000 to 1,200 W per accelerator [15, 17–19]. Under real inference workloads, GPUs typically draw 85% to 95% of their rated TDP [15]. At the node level, overhead from dual host CPUs, NVLink switches, system memory, and power-supply inefficiency adds roughly 4.5 kW, so an active 8-GPU H100 HGX node draws about 10 kW from the wall, well above the sum of the accelerators alone [15].

Table 4. Power profile of current-generation inference accelerators [15, 17–20].
PlatformTDPActive DrawIdle DrawOptimized Idle (Persistence On)
A100 SXM4 80G400 W340–380 W25–50 W5–10 W
A100 PCIe 80G300 W240–280 W25–50 W5–10 W
H100 SXM5 80G700 W600–680 W30–60 W10–12 W
H200 SXM5 141G700 W620–700 W30–60 W10–12 W
B200 SXM6 192G1,000–1,200 W900–1,100 W50–80 W15–20 W
Chart of Table 4Power draw by accelerator: active, idle, and idle with persistence mode
Active draw Idle, default drivers Idle, persistence mode on Rated TDP

Idle figures assume the paper's two cases: default drivers, and persistence mode enabled.

Source: Table 4 of the paper. The tick marks rated TDP.

View data as a table
AcceleratorTDPActiveIdleIdle, persistence on
A100 SXM4 80G400 W340 to 380 W25 to 50 W5 to 10 W
A100 PCIe 80G300 W240 to 280 W25 to 50 W5 to 10 W
H100 SXM5 80G700 W600 to 680 W30 to 60 W10 to 12 W
H200 SXM5 141G700 W620 to 700 W30 to 60 W10 to 12 W
B200 SXM6 192G1,000 to 1,200 W900 to 1,100 W50 to 80 W15 to 20 W

3.1 Cooling, Facility Overhead, and Water#

Dissipating this heat without exceeding chip junction-temperature limits requires substantial facility infrastructure, most commonly a mix of on-site chillers and evaporative cooling towers that consume water directly by evaporating it to reject heat. A mid-to-large AI data center can draw up to 5 million gallons of clean freshwater per day, comparable to the daily demand of a city of 10,000 to 50,000 people, and cooling towers evaporate roughly 70% to 80% of the water they withdraw, permanently removing it from local aquifers [21–24]. This on-site draw is compounded by off-site water use at the thermoelectric plants supplying the grid: U.S. electricity generation consumes an average of 3.1 liters of freshwater per kWh [22, 23]. Combining both effects, the average data center in the United States operates at a Water Usage Effectiveness (WUE) of 1.8 liters per kWh of IT energy, a figure established by a 2016 Lawrence Berkeley National Laboratory study and still cited as the industry baseline, though efficiency leaders report WUE below 0.2 L/kWh with liquid or immersion cooling [21–24]. At the level of a single interaction, a short 10-to-50-query chatbot session evaporates on the order of 2 liters of clean water, and generating a single 100-word email with a large model consumes roughly 500 mL [25–27].

Table 5. Cooling methodology, facility PUE, and water footprint, H100 SXM5 node [15, 18, 22–24].
Cooling MethodologyFacility PUEAnnual Energy Cost per GPU (100% Load)Hourly Cooling Water, per Node
Air-cooled racks1.30–1.50~$99213.10–15.96 L/hr
Rear-door heat exchanger / DLC1.10–1.20~$8184.93–7.39 L/hr
Fluid immersion cooling1.03–1.05~$7660.58–1.18 L/hr

At U.S. commercial electricity rates of roughly $0.12/kWh, an H100 running at half utilization costs on the order of $496 per year to power; at full utilization that figure roughly doubles, consistent with the near-linear relationship between utilization and both energy draw and its associated cooling water demand [18].

3.2 The Idling Tax#

Because LLM serving traffic is stochastic and bursty, instances must stay continuously available to meet strict TTFT service-level objectives, which makes scaling an inference server to zero impractical for most user-facing applications: reloading a model from scratch reintroduces the cold-start penalty described in Section 2. For a 70B-parameter model, that means 2 to 10 seconds of container initialization, 1 to 3 seconds of CUDA context setup, and 10 to 14 seconds transferring 70 GB of weights from local NVMe to GPU VRAM: a 15-to-45-second TTFT delay on the first query. Operators avoid this by keeping "keep-warm" nodes permanently active. Recent cloud telemetry indicates GPUs spend 14% to 76% of their operational life in this execution-idle state, waiting for a payload. An idle A100 or H100 draws 25 to 60 W under optimal conditions. In typical enterprise environments affected by persistence-driver misconfiguration, though, idle draw rises to 60 to 100 W per card (500 to 800 W per server node) unless persistence mode is explicitly enabled, in which case idle draw falls to 5 to 12 W per card [17, 20]. Multiplied across a fleet, this idling tax represents a large, non-productive consumption of energy and water whose only purpose is to keep server memory warm.

A further contributor to this thermal profile is architectural rather than operational: the historical separation of compute and memory, sometimes called the Von Neumann bottleneck, means a substantial share of the energy consumed during model execution is spent moving data between memory and compute rather than performing arithmetic. That internal data movement generates heat in its own right, which compounds the cooling and water demand described above regardless of how well a given workload is scheduled.

3.3 Embodied Carbon and Carbon-Aware Scheduling#

Sections 3.1 and 3.2 account for the water and electricity a GPU consumes while it operates. That operational figure is not the whole environmental picture. A six-year, cradle-to-grave lifecycle study of five TPU generations — covering manufacturing, transport, data-center construction, use, and end-of-life — finds that embodied emissions (manufacturing the chip, its HBM stack, and the surrounding system) account for roughly 10% to 30% of a chip's lifetime carbon footprint, with the remainder coming from operational electricity; the exact split depends heavily on whether the accounting credits an operator's carbon-free-energy purchases [63]. Table 6 reproduces the study's embodied-versus-operational breakdown across TPU generations. Embodied emissions are rising in absolute terms as chips get larger and denser: memory alone (HBM plus host DRAM) accounts for roughly 38% of manufacturing emissions on average, and manufacturing emissions per chip grew 1.8 times from the v4i to the v6e generation, driven by larger dies and more HBM [63]. Despite that increase, a Compute Carbon Intensity metric, defined as grams of CO2-equivalent per exaFLOP, improved roughly threefold from v4i to v6e, because each generation's performance gain outpaces its emissions growth [63]. The practical implication for a systems designer is the inverse of what Section 3.1's water accounting suggests in isolation: because operational electricity still dominates lifetime emissions under most accounting methods, the idling tax described in Section 3.2 remains the larger lever, but embodied carbon is not zero and grows in relative importance every year grids get cleaner, since a cleaner grid reduces operational emissions faster than manufacturing emissions fall.

Table 6. Embodied vs. operational lifetime carbon, five TPU generations, 6-year lifespan (kg CO2e) [63].
TPU GenerationTotal Embodied CO2eOperational CO2e (Market-Based / Location-Based)
v4i386 kg1,166 kg / 3,137 kg
v5e402 kg1,154 kg / 3,104 kg
v6e692 kg2,141 kg / 5,759 kg
v4693 kg2,301 kg / 6,187 kg
v5p1,101 kg4,288 kg / 11,532 kg
Chart of Table 6Lifetime carbon per TPU generation: embodied and operational
Embodied Operational (market-based)

Source: Table 6 of the paper (kg CO2e over a 6-year lifespan).

View data as a table
TPU generationEmbodied (kg)Operational, market (kg)Operational, location (kg)
v4i3861,1663,137
v5e4021,1543,104
v6e6922,1415,759
v46932,3016,187
v5p1,1014,28811,532

Operators have also begun treating carbon intensity itself as a schedulable resource rather than a fixed cost. Carbon-aware computing shifts flexible, delay-tolerant workloads in time toward hours when the local grid draws more heavily on renewable generation, and in space toward data-center regions with a higher share of carbon-free energy on the grid at that moment [65]. Google operates a production version of this idea across its full data-center fleet: a scheduler uses day-ahead grid forecasts to shift moveable compute — initially batch workloads such as video transcoding and photo processing — toward the times and locations with the highest available carbon-free-energy share, drawing on measured regional CFE scores that range from 89% in Oregon to 3% in Singapore [64]. The same mechanism does not transfer cleanly to interactive LLM inference, which is precisely the tension this paper's cold-start and idling-tax sections exist to describe: a request that must return a first token within a strict service-level objective cannot simply wait for a cleaner hour the way an overnight batch job can. Carbon-aware scheduling is consequently most applicable to the background, delay-tolerant workloads discussed in Section 4's multiplexing mechanisms, not to the latency-critical foreground path — a distinction any system attempting to combine the two has to preserve.

4. Hardware-Level GPU Multiplexing and Preemption#

When multiple execution contexts compete for a GPU, platforms share the hardware through software time-slicing, NVIDIA's Multi-Process Service (MPS), or Multi-Instance GPU (MIG) [33–35]. Time-slicing interleaves execution on a round-robin schedule, but swapping contexts in and out of GPU registers costs 10 to 100 microseconds, degrades memory bandwidth, and evicts L2 cache lines, leaving execution units idle during the transition [35, 37]. MPS instead routes kernel submissions from every client process through a centralized background server, nvidia-cuda-mps-server, which merges them into a single GPU context so that multiple client queues can overlap execution on the streaming multiprocessors (SMs) via Hyper-Q [33, 35]. Because MPS clients share that single context, however, the platform provides no hardware-level fault isolation: a fatal error in one client can trigger a GPU reset that crashes every other process sharing the context, and MPS containers typically require hostIPC: true, a meaningful operational constraint [33, 34].

Table 7. GPU sharing paradigms compared [33–35, 40].
ParadigmAllocationContext-Switch OverheadFault Isolation
Software time-slicingTemporal, coarse10–100 µsLimited (software address isolation)
NVIDIA MPSSpatial, concurrent Hyper-Q< 1 µs after startupNone — shared fault domain
NVIDIA MIGPhysical hardware slicesZero — independent channelsAbsolute hardware isolation

Administrators can bound MPS resources per client using sm_partition commands and the CUDA_MPS_SM_PARTITION environment variable, enforced through cuCtxCreate_v3 execution-affinity parameters, and clients can query their allotment via cuDeviceGetAttribute [41]. MPS performs best for small models (≤ 3B parameters) with short contexts (< 2K tokens) or prefill-heavy workloads, where it can more than double throughput [39]. That advantage decays log-linearly as context length or model size grows, because larger models and longer contexts saturate memory bandwidth during the attention-dominant decode phase and leave no idle SM slots to overlap [13, 39]. At the single-request decode stage, an H100 SXM's machine balance of roughly 295 FLOP per byte (989 TFLOPS ÷ 3.35 TB/s) means its Tensor Cores sit 99.7% idle, bound entirely by VRAM reads rather than compute [13].

4.1 Sub-Millisecond Soft Preemption#

Because NVIDIA GPUs lack proactive, low-overhead hardware preemption, a high-priority request historically had to wait for a running low-priority kernel to finish, up to 7.49 milliseconds for a dense GEMM operation [42–44]. Three recent systems close that gap in software. Hummingbird splits low-priority kernels at the thread-block level, where 99.999% of blocks complete within 390 microseconds, creating fine-grained preemption points; when a high-priority request arrives, its scheduler raises a preemption flag, halts new split-kernel launches, and schedules the urgent task, reducing average preemption delay to 121–165 microseconds while improving high-priority SLO attainment by 9.7x over spatial sharing and 3.5x over temporal sharing, with less than a 1% SLO penalty relative to running exclusively [42–44]. LithOS interposes on the CUDA API to slice grids into smaller launches on a 500-microsecond quantum, keeping physical launch overhead under 10 microseconds, and uses a transformed Amdahl's-Law model to predict task runtime and scale SM allocation via Thread Processing Cluster masking, cutting tail latency 13x relative to native MPS [36]. Valve takes a different approach, using GPU channel control to pause and resume offline background work around the online-inference lifecycle: deployed across 8,054 production GPUs, it bounds preemption to under 1 millisecond and at most once per online request, improves cluster utilization by 34.6% (a savings of roughly 2,170 GPUs), and adds under 5% TTFT and under 2% TPOT overhead, using only one line of driver modification and twenty lines of framework patching [45].

Table 8. Soft-preemption scheduling systems [36, 42–45].
SystemControl LevelTarget LatencyMechanism
HummingbirdThread-block split-kernel121–165 µsUser-space intercept; kernel-tick scheduling
LithOSInterposed kernel slicing~500 µs quantumCUDA API interposition; TPC masking
ValveChannel-controlled gating< 1,000 µsGPU channel gating; online-lifecycle tracking
Chart of Table 8GPU preemption and switching latency (log scale)

LithOS is shown by its slicing quantum, not a measured preemption latency.

Source: section 4 and Tables 7 and 8 of the paper.

View data as a table
SystemLatencyNote
Waiting for a running dense GEMM kernel7.49 msWorst case without preemption (7.49 ms)
Valve1 msBound: under 1 ms, at most once per online request
LithOS500 µsKernel slicing quantum of about 500 µs
Hummingbird121 µs to 165 µsAverage preemption delay, thread-block split kernels
Software time-slicing10 µs to 100 µsContext-switch overhead per swap

4.2 Cluster- and Orchestration-Level GPU Sharing#

Section 4.1's soft-preemption systems operate below the driver, at the level of an individual GPU's kernel scheduler. A second, largely independent layer of multiplexing happens above it, at the cluster-orchestration level, where a scheduler decides which workload gets placed on which physical GPU in the first place. Kubernetes, the dominant orchestration layer for GPU clusters, historically exposed GPUs to its scheduler only as opaque integer counts (nvidia.com/gpu: 1), which prevents the scheduler from reasoning about MIG partitions, NVLink topology, or memory headroom when placing a pod; clusters running that default device-plugin model commonly see 20% to 30% of total GPU capacity sit idle purely from placement fragmentation, not from any workload actually being finished [66]. Dynamic Resource Allocation (DRA), introduced as an alpha API in Kubernetes 1.26 and reaching beta in 1.32, replaces the opaque integer count with structured resource claims the scheduler can evaluate against real device attributes — GPU model, memory capacity, MIG-partition size, NVLink adjacency — expressed as query expressions rather than a flat quantity, which allows MIG slices and fractional GPU allocations to become first-class schedulable units instead of a workaround layered on top of whole-GPU allocation [66, 67].

Above the raw allocation API sits scheduling policy. NVIDIA's KAI Scheduler adds three capabilities the default Kubernetes scheduler lacks and that GPU batch workloads specifically need: gang scheduling, which starts every pod in a distributed training or inference job together or not at all, preventing a job from partially launching and holding GPUs idle while it waits on the rest of its pods; fair-share queuing, which enforces proportional GPU access across teams sharing a cluster instead of first-come-first-served allocation; and priority-based preemption, which lets a latency-sensitive inference workload evict a lower-priority batch job automatically rather than requiring an operator to intervene [66]. The last of these is the cluster-level analogue of Section 4.1's microsecond-scale preemption systems — the same concept of borrowing capacity from low-priority work, implemented at the granularity of whole pods and nodes rather than thread blocks and kernels. Table 9 places the two layers side by side. In practice a production deployment needs both: sub-millisecond preemption solves contention on a single already-allocated GPU, while orchestration-level bin-packing and gang scheduling determine whether that GPU was allocated efficiently across the cluster in the first place. Neither layer substitutes for the other, and a scheduler that only reasons at one level of the stack leaves the inefficiency in the other layer on the table.

Table 9. GPU multiplexing by orchestration layer [36, 42–45, 66, 67].
LayerGranularityRepresentative MechanismWhat It Solves
Driver / kernel-level (Section 4.1)Thread-block / kernelHummingbird, LithOS, ValveContention on an already-allocated GPU
Cluster-orchestration level (Section 4.2)Pod / nodeKubernetes DRA, KAI SchedulerWhether GPUs were allocated efficiently across the cluster

5. Edge Telemetry: The Privacy-Utility Duality#

To avoid the latency spikes and environmental waste of always-on cloud hosting, modern architectures increasingly lean on edge telemetry to anticipate demand before it arrives. By capturing high-frequency client-side signals such as interface focus changes, keyboard cadence, and ambient microphone level, a local telemetry client can predict an imminent LLM prompt before it is submitted, and use that prediction to pre-warm a remote GPU, stage the relevant KV context, and program MPS partitions ahead of time. The same capability, however, is dual-use: a telemetry system built to track behavior at millisecond resolution can just as easily become an apparatus of surveillance as an efficiency mechanism. Three real deployments illustrate the range of outcomes.

5.1 Windows Recall: The Surveillance-Adjacent Failure Mode#

Introduced for Copilot+ PCs with dedicated neural processing units, Windows Recall takes a screenshot of the desktop every 5 seconds, runs on-device optical character recognition on the NPU, and stores the result as a searchable, time-indexed record of the user's activity [48–51]. At launch, the resulting SQLite database was stored unencrypted, and security researchers demonstrated that a short infostealer script could exfiltrate months of keystrokes and documents in seconds, including sensitive material such as financial checkout pages and messages from apps designed to auto-delete [49, 51]. Microsoft subsequently rebuilt the feature's security architecture: Recall is now opt-in by default, its database is encrypted at rest, and the decryption key is bound to the device through the Trusted Platform Module (TPM) and gated behind Windows Hello Enhanced Sign-in Security inside a Virtualization-based Security enclave, so the database cannot be decrypted by moving the drive to another machine [48, 49]. The episode is a useful case study precisely because the underlying telemetry, a semantic index of everything visible on screen, is the most aggressive version of the capability this paper otherwise treats as an efficiency signal.

5.2 Apple Private Cloud Compute: Architecture as Privacy Guarantee#

Apple's Private Cloud Compute (PCC) takes the opposite design stance, using architecture rather than trust to bound what a cloud node can learn. Simple requests are handled entirely on-device through a local semantic index and App Intents; only requests that exceed on-device capability are routed to PCC nodes running on custom Apple silicon servers [52–55]. PCC enforces three guarantees on that path: cryptographic attestation, in which the device verifies the software image of a cloud node before connecting to it; Oblivious HTTP, which routes traffic through a third-party relay so no single party can see both the user's IP address and their request payload; and data ephemerality, in which user data is deleted immediately after a response is generated and is never retained or used for model retraining [52–54]. The residual risk Apple's own design surfaces is less about any individual mechanism and more about the user's dependence on the correctness of vendor-signed software. The whole guarantee rests on the attestation chain actually being audited and verifiable.

5.3 PRISM: Sensitivity-Aware Routing#

PRISM (Privacy-Aware Routing for Adaptive Cloud-Edge LLM Inference via Semantic Sketch Collaboration), accepted to AAAI 2026, occupies a middle position between Recall's full local capture and PCC's binary local/cloud split. An edge-side small language model profiles each prompt's entity-level sensitivity and a soft-gating module routes it down one of three paths: non-sensitive prompts go directly to the cloud LLM; highly sensitive prompts are processed entirely on-device; and moderately sensitive prompts are obfuscated with an adaptive two-layer local differential privacy mechanism before the cloud model generates a semantic sketch that the local SLM then refines into a final response using untransmitted local context [56, 57]. Across its evaluation, PRISM reduces energy consumption and latency to 40–50% of uniform or selective local-differential-privacy baselines while preserving output quality, validated with real prompts, measured energy, and heterogeneous edge-cloud model deployments [57]. Its principal failure mode is misclassification by the gating model itself: an entity-sensitivity error routes a prompt down the wrong path before any privacy mechanism has a chance to apply [56].

Table 10. Edge telemetry architectures compared [48–57].
SystemTelemetry VectorPrivacy MechanismKey Risk
Windows RecallFull-screen capture every 5 s + local OCRTPM-bound encryption, Windows Hello, opt-inInfostealer exfiltration if compromised
Apple Private Cloud ComputeOn-demand system/context embeddingsAttestation, Oblivious HTTP, data ephemeralityDepends on vendor attestation integrity
PRISMLocal named-entity sensitivity profilingAdaptive two-layer local differential privacyMisrouting by the gating model

Read together, these three systems bound the design space any predictive edge-telemetry architecture has to fit inside: telemetry rich enough to be useful for prediction, but constrained by design, not merely by policy, to prevent the Recall failure mode while avoiding a full dependence on vendor attestation.

DemoSensitivity-aware routing, step by stepIllustrative

Illustrative. The prompts, sensitivity scores and the 0.30 and 0.75 thresholds are invented for this demo. PRISM learns its gating; see section 5.3 of the paper.

Choose a prompt

1. Edge model profiles entities

Pick a prompt above.

2. Soft gate picks a path

CloudObfuscate + sketchOn-device only

3. What leaves the device

PRISM reports energy and latency at 40 to 50 percent of uniform or selective local-differential-privacy baselines (paper, section 5.3).

5.4 Privacy-Preserving ML Techniques and the Regulatory Boundary#

Sections 5.1 through 5.3 compare three deployed systems by their design choices. Underneath all three sits a smaller set of general-purpose privacy-preserving machine learning techniques, and a body of regulation that constrains how any of them can be applied to behavioral prediction specifically.

Federated learning keeps training data on-device and centralizes only model updates rather than raw records; Google's Gboard keyboard is the largest deployed example, training next-word-prediction models across more than a billion Android devices while protecting the aggregated updates with secure aggregation so that no single client's contribution is individually visible to the server [68]. Differential privacy adds calibrated statistical noise so that no individual record can be reverse-engineered from an aggregate output, parameterized by a privacy budget epsilon; Apple's on-device differential-privacy deployment, which has collected keyboard and emoji usage patterns from hundreds of millions of devices since 2016, operates with epsilon roughly between 1 and 4 per submission, where a lower value is more private and epsilon near 8 is considered only moderate protection [68]. DP-SGD extends the same idea into model training itself, clipping per-example gradients and adding Gaussian noise before any update leaves a device. On-device inference, the strongest of the three guarantees, runs the model entirely locally with no transmission at all — Apple's on-device Photos search executes CLIP-style multimodal models on the Neural Engine and builds its search index without any image ever leaving the phone [68]. These three techniques are not mutually exclusive; production-grade privacy-preserving ML systems typically combine all three, using federated learning to distribute computation, differential privacy to bound what any single update can reveal, and on-device inference to eliminate data collection for the subset of tasks that do not require the cloud at all.

None of the three is specific to behavioral-prediction telemetry — the pattern of keystroke, focus, and audio signals this paper's Sections 5.1 and 5.2 discuss as the raw input to any prediction system built on this class of data. That specific use case, predicting an upcoming user action from continuous background signal, intersects a distinct area of law. GDPR Article 22 grants a qualified right not to be subject to a decision "based solely on automated processing, including profiling," whenever that decision produces a legal effect or "similarly significantly affects" the person — a standard that extends beyond contracts and credit decisions to any automated process capable of altering a person's circumstances, behavior, or choices [69]. Where Article 22 applies, the controller owes the user meaningful information about the logic involved and a route to human review; the article does not apply where the processing is necessary to perform a contract, is authorized by law with safeguards, or rests on explicit consent [69]. A telemetry system that predicts intent and acts on that prediction — pre-warming a resource, adjusting what a user sees, routing a request — sits close enough to "profiling" that a system built for deployment in the EU has to resolve, at a design level rather than a policy level, which of Article 22's three exceptions it relies on before it ships, not after a regulator asks.

Table 11. Privacy-preserving ML techniques and their guarantees [68, 69].
TechniqueData Leaving DeviceGuarantee MechanismRepresentative Deployment
Federated learningModel updates only, not raw dataSecure aggregationGoogle Gboard (1B+ devices)
Differential privacyNoised aggregate statisticsCalibrated noise, ε ≈ 1–4Apple keyboard / emoji telemetry
On-device inferenceNothingFull local executionApple Photos on-device search

Table 11 summarizes the three techniques' guarantees, and the boundary each one draws around what a telemetry-driven prediction system built in this design space is legally and technically permitted to do.

6. Quantifying the Macro-Value of Micro-Latency#

Sections 2 and 4 establish, from published and independently measured results rather than a proposed system, that the individual stages of the cold-start and contention path can each be compressed by an order of magnitude or more: Tutti's GPU-centric KV-cache restoration cuts a 72.3-second GDS-backed restore to roughly 1.9 seconds under strict SLO constraints (Section 2.2); Hummingbird's block-level preemption reduces a multi-millisecond wait for a low-priority kernel to 121–165 microseconds (Section 4.1); Foundry's graph-template materialization removes minutes of CUDA graph capture from an autoscaling event entirely (Section 2.3). To translate that class of per-stage improvement into a fleet-scale outcome, we evaluate a representative scenario in which a shared-tier operator that adopted this combination of already-published techniques reduces user-perceived time-to-first-token by ΔT = 0.212 seconds for every query issued by a 100-million-user fleet. The figure is illustrative rather than a measured, deployed-system result: it is a representative gap consistent with the component-level latencies cited above for a typical shared-tier request under contention, not a claim about any single running system. The accounting methodology below, though, applies to any measured ΔT a production deployment observes, from any combination of the techniques this paper surveys.

6.1 Assumptions#

Active user fleet U: 100,000,000 users, averaging Q = 8 queries per day (Qtotal = 800,000,000 queries/day).

Latency savings per query: ΔT = 0.212 s, on an NVIDIA H100 SXM5 (PGPU = 0.7 kW).

Infrastructure overhead: 1.80× node overhead (host CPUs, NVLink, RAM) and 1.4× facility PUE, giving Putility = 0.7 × 1.80 × 1.4 = 1.764 kW per GPU.

Water intensity: 1.8 L/kWh on-site cooling plus 3.1 L/kWh off-site generation, giving Wrate = 4.9 L/kWh (Section 3).

Valuation: Ruser = $25.00/hour (median value of human labor time) and Rserver = $2.99/hour per H100, representative of specialist-cloud on-demand rates. 2026 market rates span roughly $1.49 to $7+/hour depending on provider tier, with a median near $3.20–$3.41/hour, so this assumption sits at the low-to-median end of the current range rather than defining it.

6.2 Time, Cost, Energy, and Water Savings#

Cumulative physical time saved per day follows directly from query volume and per-query savings:

Tsaved-day = (Qtotal × ΔT) / 3600 = (800,000,000 × 0.212) / 3600 = 47,111.11 hours/day

Scaling to weekly (×7) and monthly (×30) horizons, applying the valuation and infrastructure assumptions above, and converting energy to water through W_rate = 4.9 L/kWh yields the cumulative picture in Table 12. "Utility Energy" reflects the full P_utility = 1.764 kW/GPU figure inclusive of node overhead and cooling PUE; a pure-hardware accounting using only P_GPU = 0.7 kW yields roughly 40% of the listed energy and water figures.

Table 12. Cumulative fleet-scale savings from a representative 0.212 s per-query TTFT reduction across 100 million users.
HorizonFleet QueriesReclaimed User TimeUser TVM ($25/hr)Server TVM ($2.99/hr)Utility Energy (kWh) / Water (L)
1 Day800,000,00047,111 hours$1,177,778$140,86283,104 kWh / 407,210 L
1 Week5,600,000,000329,778 hours$8,244,444$986,036581,728 kWh / 2,850,467 L
1 Month24,000,000,0001,413,333 hours$35,333,333$4,225,8672,493,120 kWh / 12,216,288 L
Chart of Table 12Fleet savings model

Fleet queries

800,000,000

Reclaimed user time, hours

47,111

User time value

$1,177,778

Server time value

$140,862

Electricity, kWh

83,104

Water, liters

407,210

P_utility = 0.7 × 1.8 × 1.4 = 1.764 kW · W_rate = 4.9 L/kWh

Matches Table 12

The 0.212 s figure is illustrative in the paper, not a measured result.

Source: section 6 and Table 12 of the paper.

6.3 Reading the Numbers#

Two figures are worth separating conceptually even though the table combines them. User Productivity TVM values reclaimed human attention at a general labor rate; it is a proxy for the aggregate value of eliminated waiting, not a claim that every user converts idle seconds into billable work. Server Compute TVM is a harder number: it reflects GPU-seconds that a shared-tier operator no longer has to hold in a waiting state, which is directly reclaimable as either cost savings or additional served capacity. The energy and water columns follow the same utility-level accounting used throughout Section 3: they represent the electricity and cooling water a keep-warm node would have consumed over the equivalent idle window, not the marginal cost of adopting the preemption and restoration techniques themselves, which Section 4.1's sub-millisecond preemption figures and Section 2.2's 1.9-second Tutti restore time show are negligible by comparison — even summed together, they are three orders of magnitude smaller than the multi-second-to-tens-of-seconds baseline they replace. At fleet scale, a latency improvement measured in fractions of a second compounds into a monthly order of magnitude of roughly 1.4 million hours of reclaimed user time, 2.5 million kWh of avoided electricity draw, and 12 million liters of avoided freshwater consumption. Those are the same categories of physical resource quantified in Section 3, just recovered instead of spent.

7. Conclusion#

The four bottlenecks surveyed in this paper are not independent, and each has turned out to have both a hardware-adjacent layer and an orchestration layer above it that this paper's expanded treatment makes explicit. Cold-start latency (Section 2) is the reason operators keep accelerators warm, whether the remaining bottleneck is weight transfer, KV-cache restoration, or — as Section 2.3 shows — the CUDA-graph capture that recurs on every autoscaling event even after the first two are solved. Keeping accelerators warm is what drives the electrical and freshwater footprint quantified in Section 3, and Section 3.3 extends that accounting past the point a node is switched on, to the embodied carbon of manufacturing it and the carbon-aware scheduling that can shift delay-tolerant load toward cleaner grid hours — a lever unavailable to the latency-critical foreground path Section 2 describes, which is precisely why it has to be applied to the background work Section 4 discusses instead. GPU multiplexing and microsecond-scale preemption (Section 4) are the mechanism that lets a data center avoid dedicating hardware to any single tenant's worst case, and Section 4.2 shows that mechanism has a direct analogue one layer up, at cluster-orchestration granularity, where the same borrow-from-low-priority-work logic determines whether GPUs were even allocated efficiently across a fleet before any single-GPU preemption decision is made. And edge telemetry (Section 5) is the signal precise enough to tell a scheduler when action is actually needed, provided it can be collected without recreating the surveillance failure mode Windows Recall demonstrated — a design constraint Section 5.4 shows is not just a matter of good engineering taste but of a specific, applicable regulatory boundary under GDPR Article 22 once a system starts acting on predicted, rather than submitted, user intent.

None of this paper proposes a single system to resolve all four bottlenecks at once. It is deliberately the opposite: a survey of what already-published, independently measured techniques — Tutti's GPU-centric cache restoration, Hummingbird's and Valve's microsecond-scale preemption, Foundry's and ServerlessLLM's cold-start elimination, and the orchestration-level scheduling that coordinates all of them — are worth when their gains are added up and applied at fleet scale. Section 6's accounting is a reminder of why that arithmetic matters: at 100 million users, a 0.212-second improvement built from techniques that already exist in the published literature is not a UX nicety but a lever on megawatts of electricity and millions of liters of freshwater every month. In the current era of exascale AI deployment, the systems engineer optimizing a latency curve, the hardware architect accounting for a chip's embodied carbon, and the privacy engineer bounding what a telemetry pipeline is legally permitted to infer are, whether they realize it or not, working constrained versions of the same problem.

References#

All arXiv entries and the majority of vendor/technical sources below were spot-verified against their live pages during preparation of this paper; two secondary trade-press citations ([28]/[29] duplicate PDF/HTML pairs excepted) could not be re-fetched due to a rate limit and are included as originally compiled. References 61–69 were added in this revision and were fetched and spot-verified directly.

[1] Cold Start Latency In LLM Inference: Causes, Metrics & Fixes. AceCloud. https://acecloud.ai/blog/cold-start-latency-llm-inference/

[2] The GPU Cold Starts Nobody Warns You About: Autoscaling LLM Inference on Kubernetes. Medium. https://medium.com/@manikandan_t/the-gpu-cold-starts-nobody-warns-you-about-autoscaling-llm-inference-on-kubernetes-4128cb8743f1

[3] Dissecting LLM Container Cold-Start: Where the Time Actually Goes. Microsoft Tech Community. https://techcommunity.microsoft.com/blog/linuxandopensourceblog/dissecting-llm-container-cold-start-where-the-time-actually-goes/4508831

[4] GPU Cold Start on Serverless LLM Inference: 4 Fixes That Actually Work (2026). Spheron Blog. https://www.spheron.network/blog/gpu-cold-start-llm-inference-2026/

[5] Reserved GPU Capacity for AI Inference: How to Reduce Cold Starts, Latency Spikes, and Shared-Tier Jitter. GMI Cloud. https://www.gmicloud.ai/en/blog/reserved-gpu-capacity-for-ai-inference-how-to-reduce-cold-starts-latency-spikes-and-shared-tier-jitter

[6] Tutti: Making SSD-Backed KV Cache Practical for Long-Context LLM Serving. arXiv:2605.03375. https://arxiv.org/abs/2605.03375

[7] NVMe KV Cache Offloading for LLM Inference: Serve 10x More Users on the Same GPU (2026). Spheron Blog. https://www.spheron.network/blog/nvme-kv-cache-offloading-llm-inference/

[8] Tutti: Making SSD-Backed KV Cache Practical for Long-Context LLM Serving (PDF). arXiv:2605.03375. https://arxiv.org/pdf/2605.03375

[9] Deploy LMCache on GPU Cloud: Share KV Cache Across vLLM Nodes for 15x Higher Throughput (2026 Guide). Spheron Blog. https://www.spheron.network/blog/deploy-lmcache-vllm-kv-cache-sharing-gpu-cloud/

[10] When Open Source Meets Open Source: A Joint Effort Between LMCache and Mooncake. LMCache Blog. https://blog.lmcache.ai/en/2026/05/26/when-open-source-meets-open-source-a-joint-effort-between-lmcache-and-mooncake/

[11] Tutti: Making SSD-Backed KV Cache Practical for Long-Context LLM Serving. ResearchGate. https://www.researchgate.net/publication/404476613_Tutti_Making_SSD-Backed_KV_Cache_Practical_for_Long-Context_LLM_Serving

[12] LLM Inference Optimization: Techniques That Actually Reduce Latency and Cost. Runpod. https://www.runpod.io/blog/llm-inference-optimization-techniques-reduce-latency-cost

[13] GPU Inference: H100 vs A100 vs L4. Inference Engineering. https://inferenceengineering.tech/learn/gpu-inference/

[14] vLLM Custom for DGX Spark — Stream Loading and Automatic KV Cache. NVIDIA Developer Forums. https://forums.developer.nvidia.com/t/vllm-custom-for-dgx-spark-stream-loading-and-automatic-kv-cache/365798

[15] AI Inference Power Consumption and GPU Electricity Costs: 2026 Guide. Spheron Blog. https://www.spheron.network/blog/ai-inference-power-electricity-cost-2026/

[16] NVIDIA H100 Power Consumption Guide. TRG Datacenters. https://www.trgdatacenters.com/resource/nvidia-h100-power-consumption/

[17] What Is the Typical Idle Power Consumption of the NVIDIA A100 and H100 GPUs? Massed Compute. https://massedcompute.com/faq-answers/?question=What%20is%20the%20typical%20idle%20power%20consumption%20of%20the%20NVIDIA%20A100%20and%20H100%20GPUs?

[18] Estimated Annual Power Consumption Costs of NVIDIA A100 and H100 GPUs. Massed Compute. https://massedcompute.com/faq-answers/?question=What%20are%20the%20estimated%20annual%20power%20consumption%20costs%20of%20NVIDIA%20A100%20and%20H100%20GPUs%20in%20a%20typical%20data%20center?

[19] NVIDIA A100 vs. H100: Architecture and Specs. Bacloud. https://www.bacloud.com/en/blog/174/nvidia-a100-vs.-h100-architecture-and-specs.html

[20] A100 Idle Power Draw. r/homelab, Reddit. https://www.reddit.com/r/homelab/comments/1pqfafa/a100_idle_power_draw/

[21] WUE (Water Usage Effectiveness). Trane. https://www.trane.com/commercial/north-america/us/en/about-us/newsroom/glossary/water-usage-effectiveness.html

[22] Data Centers and Water Consumption. EESI. https://www.eesi.org/articles/view/data-centers-and-water-consumption

[23] What Is Water Usage Effectiveness (WUE)? Sunbird DCIM. https://www.sunbirddcim.com/glossary/water-usage-effectiveness-wue

[24] Myths vs. Reality: Data Centers and Water Usage. Florida Water and Pollution Control Operators Association. https://www.fwpcoa.org/content.aspx?page_id=5&club_id=859275&item_id=130961

[25] AI Water Footprint Calculator. Omni Calculator. https://www.omnicalculator.com/ecology/ai-water-footprint

[26] AI Water Consumption: The Thirsty Demands of AI. Brain-CA Technologies. https://brain-ca.com/ai-water-consumption-the-thirsty-demands-of-ai/

[27] AI's Hidden Water Footprint. Save the AI. https://savethe.ai/water/

[28] WattGPU: Predicting Inference Power and Latency on Unseen GPUs and LLMs. arXiv:2607.02391. https://arxiv.org/abs/2607.02391

[29] WattGPU: Predicting Inference Power and Latency on Unseen GPUs and LLMs (PDF). arXiv:2607.02391. https://arxiv.org/pdf/2607.02391

[30] Watt Counts: Energy-Aware Benchmark for Sustainable LLM Inference on Heterogeneous GPU Architectures. arXiv:2604.09048. https://arxiv.org/abs/2604.09048

[31] NVIDIA H100 Price Guide 2026: GPU Costs, Cloud Pricing & Buy vs Rent. Jarvis Labs. https://jarvislabs.ai/blog/h100-price

[32] A Guide to Data Center Water Usage Effectiveness (WUE) and Best Practices. Data Center Knowledge. https://www.datacenterknowledge.com/cooling/a-guide-to-data-center-water-usage-effectiveness-wue-and-best-practices

[33] Demystifying NVIDIA MPS: How Multi-Process Service Improves GPU Sharing and Performance. Sagar Parmar, Medium. https://sagar-parmar.medium.com/demystifying-nvidia-mps-how-multi-process-service-improves-gpu-sharing-and-performance-9f633878318a

[34] Maximize AI Infrastructure Throughput by Consolidating Underutilized GPU Workloads. NVIDIA Technical Blog. https://developer.nvidia.com/blog/maximize-ai-infrastructure-throughput-by-consolidating-underutilized-gpu-workloads/

[35] CUDA Multi-Process Service (MPS): GPU Sharing for Concurrent Workloads. Abhik Sarkar. https://www.abhik.ai/concepts/gpu-computing/cuda-mps

[36] LithOS: An Operating System for Efficient Machine Learning on GPUs. CMU CSD PhD Blog. https://www.cs.cmu.edu/~csd-phd-blog/2025/lithos/

[37] How Does CUDA Context Switching Affect Performance? Massed Compute. https://massedcompute.com/faq-answers/?question=How%20does%20CUDA%20context%20switching%20affect%20performance?

[38] Multi-Process Service. NVIDIA Documentation. https://docs.nvidia.com/deploy/mps/index.html

[39] Scaling Small LLMs with NVIDIA MPS. Databricks Blog. https://www.databricks.com/blog/scaling-small-llms-nvidia-mps

[40] Granularity- and Interference-Aware GPU Sharing with MPS. University of North Texas. https://engineering.unt.edu/cse/research/labs/csrl/files/Granularity_Alex.pdf

[41] When to Use MPS — Multi-Process Service. NVIDIA Documentation. https://docs.nvidia.com/deploy/mps/when-to-use-mps.html

[42] Hummingbird: SLO-Oriented GPU Preemption at Microsecond-Scale (PDF). arXiv:2601.04071. https://arxiv.org/pdf/2601.04071

[43] Hummingbird: SLO-Oriented GPU Preemption at Microsecond-Scale. arXiv:2601.04071. https://arxiv.org/abs/2601.04071

[44] [Literature Review] Hummingbird: SLO-Oriented GPU Preemption at Microsecond-Scale. themoonlight.io. https://www.themoonlight.io/en/review/hummingbird-slo-oriented-gpu-preemption-at-microsecond-scale

[45] Valve: Production Online–Offline Inference Colocation with Jointly-Bounded Preemption Latency and Rate. arXiv:2604.07874. https://arxiv.org/abs/2604.07874

[46] Efficient and Privacy-Aware Edge-Cloud Collaborative Inference for Large Language Models. arXiv:2607.13093. https://arxiv.org/abs/2607.13093

[47] Secure Edge AI for Data Sovereignty. Michael Hannecke, Medium. https://medium.com/@michael.hannecke/secure-edge-ai-for-data-sovereignty-e7e6c61aedd6

[48] Windows Recall Isn't the Privacy Nightmare You Think It Is. How-To Geek. https://www.howtogeek.com/windows-recall-isnt-the-privacy-nightmare-you-think-it-is/

[49] The Ongoing Controversy Surrounding Microsoft's Recall Feature. SSLs.com Blog. https://www.ssls.com/blog/the-ongoing-controversy-surrounding-microsofts-recall-feature/

[50] Phasing Out Windows Because of Recall. McNeel Forum. https://discourse.mcneel.com/t/phasing-out-windows-because-of-recall/206511

[51] Stealing Everything You've Ever Typed or Viewed on Your Own Windows PC Is Now Possible with Two Lines of Code — Inside the Copilot+ Recall Disaster. Kevin Beaumont, DoublePulsar. https://doublepulsar.com/recall-stealing-everything-youve-ever-typed-or-viewed-on-your-own-windows-pc-is-now-possible-da3e12e9465e

[52] Unlocking Apple's Private Cloud Compute: An Analysis of Privacy-Preserving Artificial Intelligence. arXiv:2605.24239. https://arxiv.org/abs/2605.24239

[53] Privacy — Features. Apple. https://www.apple.com/privacy/features/

[54] Apple Intelligence and Privacy on iPhone. Apple Support. https://support.apple.com/guide/iphone/apple-intelligence-and-privacy-iphe3f499e0e/ios

[55] Spotlight, Siri, and Apple Intelligence Privacy Explained. TWiT. https://twit.tv/posts/tech/spotlight-siri-and-apple-intelligence-privacy-explained

[56] PRISM: Privacy-Aware Routing for Adaptive Cloud–Edge LLM Inference via Semantic Sketch Collaboration. AAAI. https://ojs.aaai.org/index.php/AAAI/article/view/40041/44002

[57] PRISM: Privacy-Aware Routing for Adaptive Cloud-Edge LLM Inference via Semantic Sketch Collaboration. arXiv:2511.22788. https://arxiv.org/abs/2511.22788

[58] Latency Predictor. llm-d. https://llm-d.ai/docs/0.7/architecture/advanced/latency-predictor

[59] Predicted-Latency Based Scheduling for LLMs. llm-d. https://llm-d.ai/blog/predicted-latency-based-scheduling-for-llms

[60] Deadline-Based Scheduling for GPU with Preemption Support. ResearchGate. https://www.researchgate.net/publication/330246058_Deadline-Based_Scheduling_for_GPU_with_Preemption_Support

[61] Enabling Efficient Serverless Inference Serving for LLM (Large Language Model) in the Cloud (ServerlessLLM). arXiv:2411.15664. https://arxiv.org/abs/2411.15664

[62] Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start. arXiv:2604.06664. https://arxiv.org/html/2604.06664v1

[63] Life-Cycle Emissions of AI Hardware: A Cradle-to-Grave Approach and Generational Trends. arXiv:2502.01671. https://arxiv.org/html/2502.01671v1

[64] Google Moving Workloads Between Data Centers to Use Greener Energy. Data Center Frontier. https://www.datacenterfrontier.com/cloud/article/11428180/google-moving-workloads-between-data-centers-to-use-greener-energy

[65] Carbon-Aware Computing for Datacenters. arXiv:2106.11750. https://arxiv.org/abs/2106.11750

[66] Kubernetes GPU Orchestration in 2026: DRA, KAI Scheduler, and Grove Setup Guide. Spheron Blog. https://www.spheron.network/blog/kubernetes-gpu-orchestration-2026/

[67] Fractional GPUs and GPU Rightsizing: Stop Wasting Whole Cards. CAST AI. https://cast.ai/blog/fractional-gpu-kubernetes/

[68] Privacy-Preserving ML: Federated Learning, Differential Privacy & On-Device Inference. CalibreOS. https://www.calibreos.com/learn/mlsd-privacy-preserving-ml

[69] Artificial Intelligence, Profiling and Automated Decision Making. Baker McKenzie, Global Data and Cyber Handbook. https://resourcehub.bakermckenzie.com/en/resources/global-data-and-cyber-handbook/emea/eu/topics/artificial-intelligence-profiling-and-automated-decision-making