Skip to content
latentSource

Library

Projects, analysis, and explanations of the systems I work with.

Running Ollama Right: Local Models and Your Own Fine-Tunes on One GPU

Running Ollama Right: Local Models and Your Own Fine-Tunes on One GPU

Ollama is the least painful way to serve local models on a single GPU, but the defaults are tuned for demos, not for work. What actually matters (quantization, context memory, keep-alive) and how to get a fine-tuned model of your own running behind the same API.

Jul 1, 20266 min read
How a ReAct Agent Loop Actually Works

How a ReAct Agent Loop Actually Works

A walkthrough of one real ReAct agent run. The query was "Houston, we have a problem" and the agent returned the exact YouTube timestamp where Tom Hanks delivers the line. You can replay the recorded run right inside this article, thinking tokens and all.

Jun 1, 202611 min read
Prompt Engineering is Dead, Long Live Prompt Engineering

Prompt Engineering is Dead, Long Live Prompt Engineering

Everyone said prompt engineering was a fad. They were wrong. It just evolved. From artisanal prompting to systematic prompt design for production systems.

May 25, 20267 min read
Building AI-Powered Personal Websites

Building AI-Powered Personal Websites

The architecture of this site: bilingual articles, recorded LLM sessions, hybrid retrieval, and a read-only MCP interface. Here is what the implementation actually does.

May 18, 20262 min read
Self-Hosting AI: Running LLMs on Your Own Hardware

Self-Hosting AI: Running LLMs on Your Own Hardware

Cloud APIs are convenient but expensive. Explore how to run open-source LLMs on your own servers, from hardware selection to inference optimization.

May 4, 20266 min read
PostgreSQL as a Vector Database: pgvector in Production

PostgreSQL as a Vector Database: pgvector in Production

You don't need a separate vector database. pgvector turns PostgreSQL into a semantic search engine, with HNSW indexes, hybrid queries, and full SQL power.

Apr 27, 20267 min read
Claude Code and the Future of AI-Assisted Development

Claude Code and the Future of AI-Assisted Development

Claude Code brings an AI agent directly into your terminal. Explore what autonomous coding tools mean for software engineering workflows.

Apr 20, 20266 min read
Agentic Workflows: Orchestrating Multi-Step AI Pipelines

Agentic Workflows: Orchestrating Multi-Step AI Pipelines

Single-prompt AI is hitting its ceiling. Agentic workflows chain multiple LLM calls with tools, branching, and feedback loops to tackle complex tasks reliably.

Apr 13, 20267 min read
Knowledge Graphs Meet LLMs: Structured Reasoning at Scale

Knowledge Graphs Meet LLMs: Structured Reasoning at Scale

Hallucinations happen when models guess. Knowledge graphs give LLMs a structured, verifiable backbone, and the combination is more powerful than either alone.

Mar 16, 20265 min read
On-Device AI: The Push Toward Edge Intelligence

On-Device AI: The Push Toward Edge Intelligence

From Apple's Neural Engine to Qualcomm's NPUs, AI is moving off the cloud and onto your device. Here's what's driving the shift and what it unlocks.

Mar 9, 20265 min read
Fine-Tuning vs RAG: Choosing the Right Strategy

Fine-Tuning vs RAG: Choosing the Right Strategy

Both approaches customize LLMs for specific domains, but they solve different problems. Here's how to pick the right tool, and when to combine them.

Mar 2, 20265 min read
Speculative Decoding: The Hidden Speed Trick in Modern LLMs

Speculative Decoding: The Hidden Speed Trick in Modern LLMs

Inference is slow because tokens are generated one at a time. Speculative decoding breaks that constraint, without changing the model's output distribution.

Feb 23, 20268 min read
AI Coding Assistants: Superpower or Skill Atrophy?

AI Coding Assistants: Superpower or Skill Atrophy?

Copilot, Cursor, Claude Code, Codex. AI assistants write code faster than ever. But are developers getting better or just more dependent?

Feb 16, 20266 min read
DeepSeek R1 and the Open-Source Reasoning Revolution

DeepSeek R1 and the Open-Source Reasoning Revolution

A Chinese lab released a reasoning model that matched OpenAI o1 at a fraction of the cost, and then open-sourced it. Here's why that matters enormously.

Feb 2, 20266 min read
2025: The Year AI Went Mainstream

2025: The Year AI Went Mainstream

A look back at the breakthroughs, surprises, and pivots that defined artificial intelligence in 2025, from reasoning models to the agent revolution.

Jan 19, 20265 min read
Visualizing High-Dimensional Latent Spaces

Visualizing High-Dimensional Latent Spaces

Latent spaces encode meaning in geometry. Learn how dimensionality reduction techniques like t-SNE and UMAP make these invisible structures visible.

Jan 12, 20266 min read
Building a RAG Pipeline with Gemini

Building a RAG Pipeline with Gemini

A practical guide to building a Retrieval-Augmented Generation pipeline using Google's Gemini API, File Search Store, and embedding models.

Jan 5, 20266 min read
Allocating GPU Resources Efficiently

Allocating GPU Resources Efficiently

Optimize your ML infrastructure with strategies like MIG and quantization to cut costs without sacrificing performance.

Dec 29, 20256 min read
Building a RAG Pipeline from Scratch

Building a RAG Pipeline from Scratch

A step-by-step technical guide to constructing a Retrieval-Augmented Generation system for enterprise data.

Dec 22, 20258 min read
Global AI Adoption Rates

Global AI Adoption Rates

A visual breakdown of how AI is surging across industries, from healthcare to finance, over the next five years.

Dec 15, 20255 min read
The Cost of Intelligence

The Cost of Intelligence

Track the exponential drop in token generation costs and what this democratization means for the future of software.

Dec 8, 20255 min read
Scaling Laws of LLMs

Scaling Laws of LLMs

Understand the power laws governing AI performance: why bigger models and more compute consistently yield better results.

Dec 1, 20255 min read
Evolution of Neural Architectures

Evolution of Neural Architectures

Trace the lineage of neural networks from the humble Perceptron to the massive ResNets feeding today's AI.

Nov 24, 20256 min read
Inside the Transformer

Inside the Transformer

Peel back the layers of the Transformer architecture to visualize how attention heads process information.

Nov 17, 20256 min read
Mapping the Star Wars Universe

Mapping the Star Wars Universe

A fun data experiment: embedding fifteen Star Wars characters, testing which clusters actually show up, and finding out the model sorts them by rank and story arc instead of light and dark.

Nov 10, 20257 min read
Clustering Customer Segments

Clustering Customer Segments

Learn how K-Means clustering uncovers hidden patterns in user behavior to drive targeted business strategies.

Nov 3, 20256 min read
Visualizing Word Embeddings

Visualizing Word Embeddings

Explore the math behind language: see how 'King - Man + Woman = Queen' plays out visually in vector space.

Oct 27, 20255 min read
Navigating the Latent Space

Navigating the Latent Space

Unlock the secrets of high-dimensional data compression and how mapping complex inputs to simpler spaces drives generative AI.

Oct 20, 20255 min read
Vector Databases: The Memory Layer of Modern AI

Vector Databases: The Memory Layer of Modern AI

Embeddings are the language of semantic search, and vector databases are where they live. Learn how ANN indices power retrieval at scale.

Oct 13, 20257 min read
Chain-of-Thought Prompting: Teaching AI to Show Its Work

Chain-of-Thought Prompting: Teaching AI to Show Its Work

A single phrase, 'let's think step by step', dramatically improved LLM reasoning. Explore the science behind chain-of-thought and where it's headed.

Oct 6, 20257 min read
The Multimodal Moment: When AI Learned to See and Speak

The Multimodal Moment: When AI Learned to See and Speak

Vision-language models can now describe images, read charts, and watch videos. Here's how the architecture enables true cross-modal understanding.

Sep 29, 20255 min read
Structured Outputs: Making LLMs Reliable in Production

Structured Outputs: Making LLMs Reliable in Production

Free-form text works fine in a chat window, but production systems need JSON. A walkthrough of constrained decoding, schema enforcement, and a worked Gemini example pairing structured output with function calling.

Sep 22, 20259 min read
AI Agents: Beyond Chat to Autonomous Action

AI Agents: Beyond Chat to Autonomous Action

LLMs were just the beginning. Explore how autonomous AI agents use tools, memory, and planning to act, not just respond.

Sep 15, 20257 min read

39 of 39 articles