<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>One Layer Down on Gen-AI</title>
    <link>https://gen-ai.fyi/one-layer-down/</link>
    <description>Recent content in One Layer Down on Gen-AI</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>en</language>
    <lastBuildDate>Sat, 29 Aug 2026 00:00:00 +0000</lastBuildDate>
    <atom:link href="https://gen-ai.fyi/one-layer-down/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>My First GPU-Backed LLM Inference Run</title>
      <link>https://gen-ai.fyi/one-layer-down/first-gpu-backed-llm-run-colab/</link>
      <pubDate>Sat, 29 Aug 2026 00:00:00 +0000</pubDate>
      <guid>https://gen-ai.fyi/one-layer-down/first-gpu-backed-llm-run-colab/</guid>
      <description>I have been reading plenty about inference systems: KV cache, PagedAttention, prefill, decode, batching, and disaggregated serving. But reading about the machinery is different from running it and watching the machine do work.
This post is the first in the series that closes that gap. The goal was intentionally small: run one open-source model on a GPU, record what environment I actually got, generate 100 tokens, and capture enough numbers to make the next experiment less hand-wavy.</description>
    </item>
  </channel>
</rss>
