<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>linux on Antonio Space</title>
    <link>/tags/linux/</link>
    <description>Recent content in linux on Antonio Space</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>en</language>
    <lastBuildDate>Mon, 03 Aug 2026 20:00:00 +0800</lastBuildDate><atom:link href="/tags/linux/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Lifecycle of a Prompt on My Old PC Home Lab</title>
      <link>/posts/lifecycle-of-a-prompt-old-pc-homelab/</link>
      <pubDate>Mon, 03 Aug 2026 20:00:00 +0800</pubDate>
      
      <guid>/posts/lifecycle-of-a-prompt-old-pc-homelab/</guid>
      <description>I have spent years looking at infrastructure as a chain of small, inspectable things: a packet, a process, a file descriptor, a queue. Local LLM inference initially felt less inspectable. I would send a sentence to Ollama, the GPU fans would wake up, and text would come back. That is not a satisfying boundary for an infrastructure engineer.
So I followed one request through my old PC home lab. This is not a benchmark and it is not a claim about every Ollama release or every model.</description>
      <content>&lt;p&gt;I have spent years looking at infrastructure as a chain of small, inspectable things: a packet, a process, a file descriptor, a queue. Local LLM inference initially felt less inspectable. I would send a sentence to Ollama, the GPU fans would wake up, and text would come back. That is not a satisfying boundary for an infrastructure engineer.&lt;/p&gt;
&lt;p&gt;So I followed one request through my old PC home lab. This is not a benchmark and it is not a claim about every Ollama release or every model. It is a guided inspection of one &lt;strong&gt;measured&lt;/strong&gt; run, on this machine:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Linux host; NVIDIA RTX 3060 with 12 GB VRAM, 28 streaming multiprocessors (SMs), and compute capability 8.6.&lt;/li&gt;
&lt;li&gt;Ollama 0.20.7.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;llama3.1:8b&lt;/code&gt;, chosen for the practical reason that it fits on this GPU.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;I use three labels throughout:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Measured here&lt;/strong&gt; means I observed it on this host.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Architecture&lt;/strong&gt; means the normal job of this layer, explained in plain English.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Verify locally&lt;/strong&gt; means a tempting implementation detail that needs logs, source review, or tracing before I would state it as fact for this box.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;1-the-setup-and-request&#34;&gt;1. The setup and request&lt;/h2&gt;
&lt;p&gt;Before I sent anything, &lt;code&gt;ollama ps&lt;/code&gt; showed no loaded models. The GPU was using roughly 25 MiB, which is background driver/desktop-level usage rather than evidence that the model was resident.&lt;/p&gt;
&lt;p&gt;It is important to separate two similarly named tools before looking at the request:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;ollama ps&lt;/code&gt; is Ollama&amp;rsquo;s view of models it has loaded. Its &lt;code&gt;PROCESSOR&lt;/code&gt; column reports placement such as &lt;code&gt;100% GPU&lt;/code&gt;; its &lt;code&gt;CONTEXT&lt;/code&gt; column shows allocated context. It is the right view for &lt;em&gt;model residency&lt;/em&gt;.&lt;/li&gt;
&lt;li&gt;Linux &lt;code&gt;ps&lt;/code&gt; is the operating system&amp;rsquo;s view of processes and threads. It is the right view for &lt;em&gt;programs the OS scheduled&lt;/em&gt;. It does not know which model is loaded or how many CUDA threads are running.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Ollama documents &lt;code&gt;ollama ps&lt;/code&gt; specifically as the way to inspect loaded models and GPU/CPU placement, and its context-length guide describes &lt;code&gt;PROCESSOR&lt;/code&gt; and &lt;code&gt;CONTEXT&lt;/code&gt; as the useful fields. &lt;a href=&#34;https://docs.ollama.com/context-length&#34;&gt;Ollama: context length&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Here is the request shape from my test. The exact prompt sentence is an &lt;strong&gt;awaiting-measurement placeholder&lt;/strong&gt;: the supplied attachment that contained it is not available in this working copy, and I will not invent it. The recorded result says its rendered prompt evaluated to 22 tokens.&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; style=&#34;color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;&#34;&gt;&lt;code class=&#34;language-bash&#34; data-lang=&#34;bash&#34;&gt;&lt;span style=&#34;display:flex;&#34;&gt;&lt;span&gt;curl http://127.0.0.1:11434/api/generate &lt;span style=&#34;color:#ae81ff&#34;&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&#34;display:flex;&#34;&gt;&lt;span&gt;&lt;span style=&#34;color:#ae81ff&#34;&gt;&lt;/span&gt;  -H &lt;span style=&#34;color:#e6db74&#34;&gt;&amp;#39;Content-Type: application/json&amp;#39;&lt;/span&gt; &lt;span style=&#34;color:#ae81ff&#34;&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&#34;display:flex;&#34;&gt;&lt;span&gt;&lt;span style=&#34;color:#ae81ff&#34;&gt;&lt;/span&gt;  -d &lt;span style=&#34;color:#e6db74&#34;&gt;&amp;#39;{
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&#34;display:flex;&#34;&gt;&lt;span&gt;&lt;span style=&#34;color:#e6db74&#34;&gt;    &amp;#34;model&amp;#34;: &amp;#34;llama3.1:8b&amp;#34;,
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&#34;display:flex;&#34;&gt;&lt;span&gt;&lt;span style=&#34;color:#e6db74&#34;&gt;    &amp;#34;prompt&amp;#34;: &amp;#34;&amp;lt;paste the original 22-token test prompt here&amp;gt;&amp;#34;,
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&#34;display:flex;&#34;&gt;&lt;span&gt;&lt;span style=&#34;color:#e6db74&#34;&gt;    &amp;#34;stream&amp;#34;: false,
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&#34;display:flex;&#34;&gt;&lt;span&gt;&lt;span style=&#34;color:#e6db74&#34;&gt;    &amp;#34;keep_alive&amp;#34;: &amp;#34;15m&amp;#34;,
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&#34;display:flex;&#34;&gt;&lt;span&gt;&lt;span style=&#34;color:#e6db74&#34;&gt;    &amp;#34;options&amp;#34;: {
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&#34;display:flex;&#34;&gt;&lt;span&gt;&lt;span style=&#34;color:#e6db74&#34;&gt;      &amp;#34;num_ctx&amp;#34;: 4096,
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&#34;display:flex;&#34;&gt;&lt;span&gt;&lt;span style=&#34;color:#e6db74&#34;&gt;      &amp;#34;num_predict&amp;#34;: 128,
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&#34;display:flex;&#34;&gt;&lt;span&gt;&lt;span style=&#34;color:#e6db74&#34;&gt;      &amp;#34;temperature&amp;#34;: 0,
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&#34;display:flex;&#34;&gt;&lt;span&gt;&lt;span style=&#34;color:#e6db74&#34;&gt;      &amp;#34;seed&amp;#34;: 42
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&#34;display:flex;&#34;&gt;&lt;span&gt;&lt;span style=&#34;color:#e6db74&#34;&gt;    }
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&#34;display:flex;&#34;&gt;&lt;span&gt;&lt;span style=&#34;color:#e6db74&#34;&gt;  }&amp;#39;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The settings are intentionally boring, which makes them useful for learning:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;code&gt;model&lt;/code&gt; names the local model Ollama should schedule.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;code&gt;prompt&lt;/code&gt; is the user text. Unless &lt;code&gt;raw&lt;/code&gt; is requested, it is normally rendered through the model&amp;rsquo;s template before inference.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;code&gt;stream: false&lt;/code&gt; asks for one final JSON response instead of partial response chunks. The model can still generate incrementally; the server accumulates the chunks before responding. This is visible in the &lt;a href=&#34;https://github.com/ollama/ollama/blob/v0.20.7/server/routes.go#L591-L629&#34;&gt;0.20.7 generate handler&lt;/a&gt;, and documented in the &lt;a href=&#34;https://docs.ollama.com/api/generate&#34;&gt;Generate API&lt;/a&gt;.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;code&gt;keep_alive: &amp;quot;15m&amp;quot;&lt;/code&gt; asks Ollama to retain the model in memory for 15 minutes after the request, making a near-term follow-up eligible to avoid a cold load. The API value overrides the server default. &lt;a href=&#34;https://docs.ollama.com/faq#how-do-i-keep-a-model-loaded-in-memory-or-make-it-unload-immediately&#34;&gt;Ollama FAQ: keeping models loaded&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;code&gt;num_ctx: 4096&lt;/code&gt; reserves a 4,096-token context window: the maximum token history the model can access for this request. More context costs more memory. &lt;a href=&#34;https://docs.ollama.com/context-length&#34;&gt;Ollama: context length&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;code&gt;num_predict: 128&lt;/code&gt; caps the requested generated-token budget at 128; the model may stop earlier.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;code&gt;temperature: 0&lt;/code&gt; asks for non-random, greedy-style selection rather than deliberately varied sampling.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;code&gt;seed: 42&lt;/code&gt; fixes the sampler seed. Together with temperature zero, it makes comparison easier, but it is not a universal promise of byte-for-byte reproducibility across different builds, devices, or future releases.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The public API calls &lt;code&gt;total_duration&lt;/code&gt;, &lt;code&gt;load_duration&lt;/code&gt;, prompt-evaluation counts/duration, and output-evaluation counts/duration out explicitly. &lt;a href=&#34;https://docs.ollama.com/api/generate#response&#34;&gt;Ollama: Generate API response fields&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;For this cold request, I recorded:&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; style=&#34;color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;&#34;&gt;&lt;code class=&#34;language-text&#34; data-lang=&#34;text&#34;&gt;&lt;span style=&#34;display:flex;&#34;&gt;&lt;span&gt;total duration                  60.973 s
&lt;/span&gt;&lt;/span&gt;&lt;span style=&#34;display:flex;&#34;&gt;&lt;span&gt;loading / preparation           59.335 s
&lt;/span&gt;&lt;/span&gt;&lt;span style=&#34;display:flex;&#34;&gt;&lt;span&gt;prompt evaluation               22 tokens in 60.1 ms
&lt;/span&gt;&lt;/span&gt;&lt;span style=&#34;display:flex;&#34;&gt;&lt;span&gt;generation                      76 tokens in 1.259 s  ≈ 60.4 tokens/s
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;That first number is end-to-end service time. The 59.335 seconds is Ollama&amp;rsquo;s load/preparation metric, not a stopwatch around disk reads alone. It can include scheduling, starting or connecting a runner, allocating memory, backend/model initialization, and transfers as well as storage I/O. I would need tracing to apportion it further.&lt;/p&gt;
&lt;h2 id=&#34;2-ollamas-server-runner-linux-processes-and-threads&#34;&gt;2. Ollama&amp;rsquo;s server, runner, Linux processes, and threads&lt;/h2&gt;
&lt;p&gt;The first useful correction to my mental model was that “Ollama” is not one undifferentiated thing.&lt;/p&gt;
&lt;p&gt;&lt;img src=&#34;/images/prompt-lifecycle-inference-stack.png&#34; alt=&#34;Hand-drawn stack from prompt through Ollama, CUDA, and the RTX 3060&#34;&gt;&lt;/p&gt;
&lt;p&gt;For this sample, Linux showed:&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; style=&#34;color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;&#34;&gt;&lt;code class=&#34;language-text&#34; data-lang=&#34;text&#34;&gt;&lt;span style=&#34;display:flex;&#34;&gt;&lt;span&gt;ollama server  PID 2010
&lt;/span&gt;&lt;/span&gt;&lt;span style=&#34;display:flex;&#34;&gt;&lt;span&gt;└─ ollama runner  PID 1966421  (parent PID 2010)
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Both processes had 11 &lt;strong&gt;Linux threads&lt;/strong&gt; at the instant I sampled them. A Linux thread is an OS scheduling unit inside a process: for example, it may run Go runtime work, HTTP handling, polling, or native backend code.&lt;/p&gt;
&lt;p&gt;That count is &lt;em&gt;not&lt;/em&gt; the count of CUDA threads. CUDA threads are tiny GPU execution lanes created by a kernel launch; they are grouped into blocks and scheduled on SMs. Nor is a CUDA block a Linux process. They sit on completely different sides of the driver boundary.&lt;/p&gt;
&lt;p&gt;At a high level, the server accepts and validates the HTTP request, chooses or reuses a runner, renders the request, and sends back HTTP. The runner is the inference-serving child that performs model execution through the selected backend. The version-tagged source shows the server consolidating model/request options, asking the scheduler for a runner, then invoking completion on it. &lt;a href=&#34;https://github.com/ollama/ollama/blob/v0.20.7/server/routes.go#L120-L170&#34;&gt;Ollama 0.20.7: scheduling and options&lt;/a&gt; and &lt;a href=&#34;https://github.com/ollama/ollama/blob/v0.20.7/server/routes.go#L510-L590&#34;&gt;completion call&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;The process tree proves that a child runner existed with parent 2010. It does &lt;strong&gt;not&lt;/strong&gt; tell me whether that child was created with &lt;code&gt;fork()&lt;/code&gt;, &lt;code&gt;clone()&lt;/code&gt;, &lt;code&gt;posix_spawn()&lt;/code&gt;, or another mechanism. &lt;code&gt;strace -f&lt;/code&gt; or audit data is needed for that claim; I do not make it here.&lt;/p&gt;
&lt;h2 id=&#34;3-loading-model-weights-from-storage-into-ram-and-vram&#34;&gt;3. Loading model weights from storage into RAM and VRAM&lt;/h2&gt;
&lt;p&gt;After the request, &lt;code&gt;ollama ps&lt;/code&gt; reported the model as &lt;code&gt;100% GPU&lt;/code&gt; with a 4,096-token context. It reported 5,374 MiB allocated to the runner. That is strong evidence that this configuration fit and was fully offloaded, not a claim that every byte in the process or every GPU allocation is model weight.&lt;/p&gt;
&lt;p&gt;The useful physical picture is:&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; style=&#34;color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;&#34;&gt;&lt;code class=&#34;language-text&#34; data-lang=&#34;text&#34;&gt;&lt;span style=&#34;display:flex;&#34;&gt;&lt;span&gt;model files on storage  →  host RAM / OS page cache  →  GPU VRAM
&lt;/span&gt;&lt;/span&gt;&lt;span style=&#34;display:flex;&#34;&gt;&lt;span&gt;                                      ↘ metadata, CPU-side work
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Model &lt;strong&gt;weights&lt;/strong&gt; are the learned numeric arrays that dominate the static footprint. The runner must make them available to its inference backend. On a fully GPU-offloaded run, the relevant tensors are placed in GPU memory so the GPU can use them without repeatedly crossing PCIe for each operation. The runner also needs working memory for the execution graph and a KV cache, which I will meet shortly.&lt;/p&gt;
&lt;p&gt;The exact balance between file mapping, reads, host buffers, driver allocations, and transfers is &lt;strong&gt;verify locally&lt;/strong&gt; territory. The observation I have is “59.335 seconds of loading/preparation,” not a disk-throughput measurement. In particular, I cannot honestly turn that duration into “the SSD took 59 seconds.”&lt;/p&gt;
&lt;h2 id=&#34;4-templates-tokenization-embeddings-and-prefill&#34;&gt;4. Templates, tokenization, embeddings, and prefill&lt;/h2&gt;
&lt;p&gt;The plain prompt string is not immediately a GPU calculation.&lt;/p&gt;
&lt;p&gt;First, Ollama normally applies the model&amp;rsquo;s prompt template. A template adds the model-specific conversation markers and indicates that it is the assistant&amp;rsquo;s turn to answer. In 0.20.7, the generate handler uses the model template unless &lt;code&gt;raw&lt;/code&gt; is set, constructs a user message, and renders a prompt before calling completion. &lt;a href=&#34;https://github.com/ollama/ollama/blob/v0.20.7/server/routes.go#L410-L483&#34;&gt;Ollama 0.20.7: template path&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Then the tokenizer converts that rendered text into integer token IDs. A token is a piece of text, not reliably one word: it might be a word, punctuation, or part of a word. The 22 measured prompt tokens count this model-facing representation; it need not equal the number of whitespace-separated words in my original sentence.&lt;/p&gt;
&lt;p&gt;Each token ID is looked up in an &lt;strong&gt;embedding&lt;/strong&gt; table, turning the integer into a vector of numbers the network can process. The runner evaluates the prompt tokens together in a phase usually called &lt;strong&gt;prefill&lt;/strong&gt;. Prefill builds the initial attention state and stores the useful past information in the key/value cache. Here, 22 prompt tokens took 60.1 ms.&lt;/p&gt;
&lt;p&gt;“Together” is deliberately not “one kernel.” An inference backend builds a computation graph and may launch many kernels—matrix operations, normalisation, attention-related work, copies, and more. The exact launch sequence depends on model, backend, build, shapes, and runtime choices. A GPU profiler is what would establish the sequence on this host.&lt;/p&gt;
&lt;h2 id=&#34;5-cuda-contexts-kernels-grids-blocks-warps-sms-and-memory&#34;&gt;5. CUDA contexts, kernels, grids, blocks, warps, SMs, and memory&lt;/h2&gt;
&lt;p&gt;This is where the vocabulary gets dense, so I prefer a small map over pretending all of it is the same thing.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A &lt;strong&gt;CUDA context&lt;/strong&gt; is driver-managed state that lets a host process use a GPU: allocations, modules, streams, and more live in that world. I expect the CUDA backend used by the runner to establish/use a context, but the number and timing of contexts on this installation requires verification.&lt;/li&gt;
&lt;li&gt;A &lt;strong&gt;kernel&lt;/strong&gt; is a function launched from CPU-side code to execute on the GPU. One high-level model operation can require multiple kernels; a kernel can also implement work from more than one conceptual step.&lt;/li&gt;
&lt;li&gt;A kernel launch defines a &lt;strong&gt;grid&lt;/strong&gt; of &lt;strong&gt;blocks&lt;/strong&gt;. A block is a cooperating group of CUDA threads, not a process.&lt;/li&gt;
&lt;li&gt;NVIDIA executes threads in &lt;strong&gt;warps&lt;/strong&gt;, usually 32 threads travelling through instructions together. Again: these are GPU lanes, unrelated to the 11 Linux threads I saw.&lt;/li&gt;
&lt;li&gt;An &lt;strong&gt;SM&lt;/strong&gt; (streaming multiprocessor) is a hardware worker on the GPU that schedules and executes blocks/warps. This RTX 3060 has 28 SMs.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;For inference, the runner asks CUDA software to launch kernels. The CUDA driver sends that work to the RTX 3060; the GPU schedules blocks onto its SMs and reads/writes VRAM. Global VRAM holds large data such as weights and cache buffers. The GPU also has much smaller, faster on-chip memories and registers, which kernels use according to their implementation.&lt;/p&gt;
&lt;p&gt;Compute capability 8.6 identifies the GPU architecture target exposed by CUDA. It is useful compatibility context, not a performance result on its own.&lt;/p&gt;
&lt;h2 id=&#34;6-kv-cache-iterative-decoding-token-selection-and-returning-text&#34;&gt;6. KV cache, iterative decoding, token selection, and returning text&lt;/h2&gt;
&lt;p&gt;Attention lets a transformer use earlier tokens while processing the current one. Recomputing the full history from scratch for every output token would be wasteful. The &lt;strong&gt;KV cache&lt;/strong&gt; saves the “keys” and “values” produced for previous tokens so later steps can reuse them.&lt;/p&gt;
&lt;p&gt;After prefill, generation becomes iterative:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;The backend evaluates the current sequence state and produces scores—logits—for possible next tokens.&lt;/li&gt;
&lt;li&gt;The sampling policy selects one token. With the settings above, temperature zero and seed 42 keep this test intentionally restrained.&lt;/li&gt;
&lt;li&gt;The selected token is appended to the sequence and its key/value state is added to the cache.&lt;/li&gt;
&lt;li&gt;The process repeats until the model emits a stop condition or reaches the output limit.&lt;/li&gt;
&lt;li&gt;Token IDs are detokenized into text. The runner returns response chunks to the server; with &lt;code&gt;stream: false&lt;/code&gt;, the server collects them into one JSON response.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The measured output was 76 tokens in 1.259 seconds, about 60.4 tokens per second. That rate describes this one response&amp;rsquo;s decode interval. It does not include the cold-load time and should not be advertised as a general RTX 3060 benchmark.&lt;/p&gt;
&lt;p&gt;The cache also explains a subtle warm-path possibility: repeated prompts or conversations with a shared prefix &lt;strong&gt;may&lt;/strong&gt; reuse cached prompt state, reducing the prompt work. Whether a particular request qualifies, and how much state 0.20.7 reused in this setup, needs a controlled repeated-request capture. Newer API documentation exposes &lt;code&gt;prompt_eval_cached_count&lt;/code&gt;, a useful field when available. &lt;a href=&#34;https://docs.ollama.com/api/generate&#34;&gt;Ollama: Generate API&lt;/a&gt;&lt;/p&gt;
&lt;h2 id=&#34;7-cold-versus-warm-requests-and-what-remains-afterward&#34;&gt;7. Cold versus warm requests, and what remains afterward&lt;/h2&gt;
&lt;p&gt;My first request was cold because no model was loaded before it. Its 60.973-second total was overwhelmingly preparation/loading time; prompt evaluation and 76-token generation together were roughly 1.319 seconds.&lt;/p&gt;
&lt;p&gt;With &lt;code&gt;keep_alive: &amp;quot;15m&amp;quot;&lt;/code&gt;, the expected next state is not “everything disappears when HTTP returns.” The runner/model can remain resident and visible in &lt;code&gt;ollama ps&lt;/code&gt; until the keep-alive expiry or another scheduling decision unloads it. On this test, that resident state was:&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; style=&#34;color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;&#34;&gt;&lt;code class=&#34;language-text&#34; data-lang=&#34;text&#34;&gt;&lt;span style=&#34;display:flex;&#34;&gt;&lt;span&gt;PROCESSOR: 100% GPU
&lt;/span&gt;&lt;/span&gt;&lt;span style=&#34;display:flex;&#34;&gt;&lt;span&gt;CONTEXT:   4096 tokens
&lt;/span&gt;&lt;/span&gt;&lt;span style=&#34;display:flex;&#34;&gt;&lt;span&gt;runner allocation observed: 5,374 MiB
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;A warm request can skip the costly model-load path while the model is resident. It still has to accept HTTP, render/tokenize the prompt, run prefill for uncached input, decode tokens, and return the response. A warm request is therefore “less work,” not “no work.” After 15 minutes of idleness, the model may be unloaded and the next request may again pay a cold-start cost. Ollama documents five minutes as the default and allows each API request to override it with &lt;code&gt;keep_alive&lt;/code&gt;. &lt;a href=&#34;https://docs.ollama.com/faq#how-do-i-keep-a-model-loaded-in-memory-or-make-it-unload-immediately&#34;&gt;Ollama FAQ: model residency&lt;/a&gt;&lt;/p&gt;
&lt;h2 id=&#34;what-i-would-measure-next&#34;&gt;What I would measure next&lt;/h2&gt;
&lt;p&gt;The following commands add evidence without pretending the current observations are more precise than they are. Run them on the Linux host while the model is idle, during a cold request, and during a warm repeat. They are read-only except for the intentionally sent test request.&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; style=&#34;color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;&#34;&gt;&lt;code class=&#34;language-bash&#34; data-lang=&#34;bash&#34;&gt;&lt;span style=&#34;display:flex;&#34;&gt;&lt;span&gt;&lt;span style=&#34;color:#75715e&#34;&gt;# Ollama&amp;#39;s view: loaded model, offload split, context, expiry&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style=&#34;display:flex;&#34;&gt;&lt;span&gt;ollama ps
&lt;/span&gt;&lt;/span&gt;&lt;span style=&#34;display:flex;&#34;&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style=&#34;display:flex;&#34;&gt;&lt;span&gt;&lt;span style=&#34;color:#75715e&#34;&gt;# Linux process tree and Linux thread counts (not CUDA-thread counts)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style=&#34;display:flex;&#34;&gt;&lt;span&gt;ps -o pid,ppid,nlwp,stat,cmd -p 2010,1966421
&lt;/span&gt;&lt;/span&gt;&lt;span style=&#34;display:flex;&#34;&gt;&lt;span&gt;ps -T -p &lt;span style=&#34;color:#ae81ff&#34;&gt;2010&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style=&#34;display:flex;&#34;&gt;&lt;span&gt;ps -T -p &lt;span style=&#34;color:#ae81ff&#34;&gt;1966421&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style=&#34;display:flex;&#34;&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style=&#34;display:flex;&#34;&gt;&lt;span&gt;&lt;span style=&#34;color:#75715e&#34;&gt;# GPU allocations and processes; sample it while the request runs&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style=&#34;display:flex;&#34;&gt;&lt;span&gt;nvidia-smi --query-gpu&lt;span style=&#34;color:#f92672&#34;&gt;=&lt;/span&gt;name,memory.used,memory.total --format&lt;span style=&#34;color:#f92672&#34;&gt;=&lt;/span&gt;csv,noheader
&lt;/span&gt;&lt;/span&gt;&lt;span style=&#34;display:flex;&#34;&gt;&lt;span&gt;nvidia-smi pmon -c &lt;span style=&#34;color:#ae81ff&#34;&gt;1&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style=&#34;display:flex;&#34;&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style=&#34;display:flex;&#34;&gt;&lt;span&gt;&lt;span style=&#34;color:#75715e&#34;&gt;# Record the full JSON timing fields from the same request shape&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style=&#34;display:flex;&#34;&gt;&lt;span&gt;curl -s http://127.0.0.1:11434/api/generate &lt;span style=&#34;color:#ae81ff&#34;&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&#34;display:flex;&#34;&gt;&lt;span&gt;&lt;span style=&#34;color:#ae81ff&#34;&gt;&lt;/span&gt;  -H &lt;span style=&#34;color:#e6db74&#34;&gt;&amp;#39;Content-Type: application/json&amp;#39;&lt;/span&gt; &lt;span style=&#34;color:#ae81ff&#34;&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&#34;display:flex;&#34;&gt;&lt;span&gt;&lt;span style=&#34;color:#ae81ff&#34;&gt;&lt;/span&gt;  -d @request.json | jq &lt;span style=&#34;color:#e6db74&#34;&gt;&amp;#39;{total_duration,load_duration,prompt_eval_count,prompt_eval_duration,eval_count,eval_duration}&amp;#39;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style=&#34;display:flex;&#34;&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style=&#34;display:flex;&#34;&gt;&lt;span&gt;&lt;span style=&#34;color:#75715e&#34;&gt;# Establish process-creation syscalls only if needed. This is intrusive enough&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style=&#34;display:flex;&#34;&gt;&lt;span&gt;&lt;span style=&#34;color:#75715e&#34;&gt;# to use in a short, isolated test window.&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style=&#34;display:flex;&#34;&gt;&lt;span&gt;sudo strace -ff -e trace&lt;span style=&#34;color:#f92672&#34;&gt;=&lt;/span&gt;process -p &lt;span style=&#34;color:#ae81ff&#34;&gt;2010&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;For CUDA context and kernel-level evidence, I would capture an Nsight Systems trace in a controlled run, then inspect launches, streams, memory copies, and CPU/GPU overlap. That is the right tool to answer “which kernels actually ran?” It is not something I can infer from &lt;code&gt;ps&lt;/code&gt;, &lt;code&gt;ollama ps&lt;/code&gt;, or a duration field.&lt;/p&gt;
&lt;p&gt;Near the bottom of this investigation, the whole journey finally looks ordinary again: a request crosses a server boundary, a scheduled worker executes a graph, a driver submits work to hardware, and measured timings tell me where to look next.&lt;/p&gt;
&lt;p&gt;&lt;img src=&#34;/images/prompt-lifecycle-request-to-response.png&#34; alt=&#34;Hand-drawn lifecycle: HTTP request, model loading, tokenization, prefill, decoding, and HTTP response&#34;&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Version note: the observations are from Ollama 0.20.7. The source links above are pinned to that tag where implementation behavior is discussed; public documentation can describe newer releases. Re-run the measurements after upgrading.&lt;/em&gt;&lt;/p&gt;
&lt;h2 id=&#34;watch-the-short-version&#34;&gt;Watch the short version&lt;/h2&gt;
&lt;div style=&#34;position: relative; width: 100%; max-width: 360px; aspect-ratio: 9 / 16; margin: 1.5rem auto;&#34;&gt;
  &lt;iframe
    src=&#34;https://www.youtube.com/embed/9rK06fNZPos&#34;
    title=&#34;Lifecycle of a Prompt on My Old PC Home Lab&#34;
    style=&#34;position: absolute; inset: 0; width: 100%; height: 100%; border: 0;&#34;
    allow=&#34;accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share&#34;
    allowfullscreen&gt;
  &lt;/iframe&gt;
&lt;/div&gt;
</content>
    </item>
    
  </channel>
</rss>
