Skip to content
latentSource
Machine learning · AI · Data science

Notes from buildingwith data and AI.

I'm Oswaldo Orona. I write about the systems I build, from retrieval and AI agents to the data work underneath. Read the decisions, explore the results, and inspect the code.

Agents

Systems I built, with code and decisions you can inspect.

JobMap
Agent

JobMap

Describe the work you want to do, explore matching job postings, and inspect the evidence behind Jev's relevance judgments.

Explore the agent →
Star Wars embedding atlas
Agent

Star Wars embedding atlas

An interactive explorer for fifteen character descriptions, with similarity scores, projection views, and a way to test what changes when you remove their named relationships.

Explore the agent →
ReActReel
Agent

ReActReel

A movie and music search agent that turns a remembered line into a playable video, with the tool calls visible beside the result.

Explore the agent →
Masking before the agent handoff
Agent

Masking before the agent handoff

A database transfer workflow that masks customer fields before rows reach an agent or its mailbox.

Explore the agent →
Mexico City transit assistant
Agent

Mexico City transit assistant

A conversational transit map that resolves places and asks a routing engine to calculate the journey.

Explore the agent →
Memo Ocho Bits
Agent

Memo Ocho Bits

An AI participant in a World Cup prediction pool, with saved forecasts and a leaderboard scored alongside human players.

Explore the agent →
Investigating the Gulf oil spill
Agent

Investigating the Gulf oil spill

A geospatial investigation built to challenge the Mexican government’s account of the oil spill using vessel records, satellite layers, and dated statements.

Explore the agent →
The blog as an MCP server
Agent

The blog as an MCP server

A public MCP endpoint that lets an assistant search the blog, browse its catalog, and retrieve complete articles.

Explore the agent →
About my background

Library

Projects, analysis, and explanations of the systems I work with.

Jev ranked jobs better. MiniLM found about as many strong matches.

Jev ranked jobs better. MiniLM found about as many strong matches.

I compared Jev with three local embedding models on 237 job postings. Jev led on graded relevance, but a stricter definition of a good match changed the conclusion.

AI evaluationSemantic searchJev+1
Oct 1, 20268 min read
I Built a Transit Assistant That Has to Calculate the Route

I Built a Transit Assistant That Has to Calculate the Route

A Gemini agent connects to MCP tools backed by PostGIS and pgRouting. The interesting part is resolving a place, choosing stations, and showing where the answer came from.

Sep 21, 20267 min read
What a Moving Dot Can Hide: Building an Oil-Spill Data Explorer

What a Moving Dot Can Hide: Building an Oil-Spill Data Explorer

I combined vessel records, satellite layers, and a timeline in a geospatial app. Then I checked which parts of the animation were observations and which were interpolation.

Sep 21, 20266 min read
Turning This Blog Into an MCP Server

Turning This Blog Into an MCP Server

I exposed the blog’s retrieval system as three read-only MCP tools. The implementation is small enough to inspect, and the same search tool appears in recorded agent sessions.

Jul 5, 20268 min read
Running Ollama Right: Local Models and Your Own Fine-Tunes on One GPU

Running Ollama Right: Local Models and Your Own Fine-Tunes on One GPU

Ollama is the least painful way to serve local models on a single GPU, but the defaults are tuned for demos, not for work. What actually matters (quantization, context memory, keep-alive) and how to get a fine-tuned model of your own running behind the same API.

Jul 1, 20266 min read
I Put an AI in My World Cup Pool. Here Is What I Can Measure.

I Put an AI in My World Cup Pool. Here Is What I Can Measure.

A grounded prediction agent competed under the same scoring rules as people. The published group-stage chart reports 47 correct outcomes from 72 picks; here is the implementation and what that result still needs to establish.

Jun 28, 20267 min read
Two Pi Agents and a Customer Table: Masking Before the Handoff

Two Pi Agents and a Customer Table: Masking Before the Handoff

I put masking inside the producer’s data tool, then limited the tools that agent could call. Here is the code, the synthetic example, and the boundary this proof of concept actually provides.

Jun 8, 20266 min read
How a ReAct Agent Loop Actually Works

How a ReAct Agent Loop Actually Works

A walkthrough of one real ReAct agent run. The query was "Houston, we have a problem" and the agent returned the exact YouTube timestamp where Tom Hanks delivers the line. You can replay the recorded run right inside this article, thinking tokens and all.

Jun 1, 202611 min read
Prompt Engineering is Dead, Long Live Prompt Engineering

Prompt Engineering is Dead, Long Live Prompt Engineering

Everyone said prompt engineering was a fad. They were wrong. It just evolved. From artisanal prompting to systematic prompt design for production systems.

May 25, 20267 min read
Building AI-Powered Personal Websites

Building AI-Powered Personal Websites

The architecture of this site: bilingual articles, recorded LLM sessions, hybrid retrieval, and a read-only MCP interface. Here is what the implementation actually does.

May 18, 20262 min read
Self-Hosting AI: Running LLMs on Your Own Hardware

Self-Hosting AI: Running LLMs on Your Own Hardware

Cloud APIs are convenient but expensive. Explore how to run open-source LLMs on your own servers, from hardware selection to inference optimization.

May 4, 20266 min read
PostgreSQL as a Vector Database: pgvector in Production

PostgreSQL as a Vector Database: pgvector in Production

You don't need a separate vector database. pgvector turns PostgreSQL into a semantic search engine, with HNSW indexes, hybrid queries, and full SQL power.

Apr 27, 20267 min read
Claude Code and the Future of AI-Assisted Development

Claude Code and the Future of AI-Assisted Development

Claude Code brings an AI agent directly into your terminal. Explore what autonomous coding tools mean for software engineering workflows.

Apr 20, 20266 min read
Agentic Workflows: Orchestrating Multi-Step AI Pipelines

Agentic Workflows: Orchestrating Multi-Step AI Pipelines

Single-prompt AI is hitting its ceiling. Agentic workflows chain multiple LLM calls with tools, branching, and feedback loops to tackle complex tasks reliably.

Apr 13, 20267 min read
Knowledge Graphs Meet LLMs: Structured Reasoning at Scale

Knowledge Graphs Meet LLMs: Structured Reasoning at Scale

Hallucinations happen when models guess. Knowledge graphs give LLMs a structured, verifiable backbone, and the combination is more powerful than either alone.

Mar 16, 20265 min read
On-Device AI: The Push Toward Edge Intelligence

On-Device AI: The Push Toward Edge Intelligence

From Apple's Neural Engine to Qualcomm's NPUs, AI is moving off the cloud and onto your device. Here's what's driving the shift and what it unlocks.

Mar 9, 20265 min read
Fine-Tuning vs RAG: Choosing the Right Strategy

Fine-Tuning vs RAG: Choosing the Right Strategy

Both approaches customize LLMs for specific domains, but they solve different problems. Here's how to pick the right tool, and when to combine them.

Mar 2, 20265 min read
Speculative Decoding: The Hidden Speed Trick in Modern LLMs

Speculative Decoding: The Hidden Speed Trick in Modern LLMs

Inference is slow because tokens are generated one at a time. Speculative decoding breaks that constraint, without changing the model's output distribution.

Feb 23, 20268 min read
AI Coding Assistants: Superpower or Skill Atrophy?

AI Coding Assistants: Superpower or Skill Atrophy?

Copilot, Cursor, Claude Code, Codex. AI assistants write code faster than ever. But are developers getting better or just more dependent?

Feb 16, 20266 min read
DeepSeek R1 and the Open-Source Reasoning Revolution

DeepSeek R1 and the Open-Source Reasoning Revolution

A Chinese lab released a reasoning model that matched OpenAI o1 at a fraction of the cost, and then open-sourced it. Here's why that matters enormously.

Feb 2, 20266 min read
2025: The Year AI Went Mainstream

2025: The Year AI Went Mainstream

A look back at the breakthroughs, surprises, and pivots that defined artificial intelligence in 2025, from reasoning models to the agent revolution.

Jan 19, 20265 min read
Visualizing High-Dimensional Latent Spaces

Visualizing High-Dimensional Latent Spaces

Latent spaces encode meaning in geometry. Learn how dimensionality reduction techniques like t-SNE and UMAP make these invisible structures visible.

Jan 12, 20266 min read
Building a RAG Pipeline with Gemini

Building a RAG Pipeline with Gemini

A practical guide to building a Retrieval-Augmented Generation pipeline using Google's Gemini API, File Search Store, and embedding models.

Jan 5, 20266 min read
Allocating GPU Resources Efficiently

Allocating GPU Resources Efficiently

Optimize your ML infrastructure with strategies like MIG and quantization to cut costs without sacrificing performance.

Dec 29, 20256 min read
Building a RAG Pipeline from Scratch

Building a RAG Pipeline from Scratch

A step-by-step technical guide to constructing a Retrieval-Augmented Generation system for enterprise data.

Dec 22, 20258 min read
Global AI Adoption Rates

Global AI Adoption Rates

A visual breakdown of how AI is surging across industries, from healthcare to finance, over the next five years.

Dec 15, 20255 min read
The Cost of Intelligence

The Cost of Intelligence

Track the exponential drop in token generation costs and what this democratization means for the future of software.

Dec 8, 20255 min read
Scaling Laws of LLMs

Scaling Laws of LLMs

Understand the power laws governing AI performance: why bigger models and more compute consistently yield better results.

Dec 1, 20255 min read
Evolution of Neural Architectures

Evolution of Neural Architectures

Trace the lineage of neural networks from the humble Perceptron to the massive ResNets feeding today's AI.

Nov 24, 20256 min read
Inside the Transformer

Inside the Transformer

Peel back the layers of the Transformer architecture to visualize how attention heads process information.

Nov 17, 20256 min read
Mapping the Star Wars Universe

Mapping the Star Wars Universe

A fun data experiment: embedding fifteen Star Wars characters, testing which clusters actually show up, and finding out the model sorts them by rank and story arc instead of light and dark.

Nov 10, 20257 min read
Clustering Customer Segments

Clustering Customer Segments

Learn how K-Means clustering uncovers hidden patterns in user behavior to drive targeted business strategies.

Nov 3, 20256 min read
Visualizing Word Embeddings

Visualizing Word Embeddings

Explore the math behind language: see how 'King - Man + Woman = Queen' plays out visually in vector space.

Oct 27, 20255 min read
Navigating the Latent Space

Navigating the Latent Space

Unlock the secrets of high-dimensional data compression and how mapping complex inputs to simpler spaces drives generative AI.

Oct 20, 20255 min read
Vector Databases: The Memory Layer of Modern AI

Vector Databases: The Memory Layer of Modern AI

Embeddings are the language of semantic search, and vector databases are where they live. Learn how ANN indices power retrieval at scale.

Oct 13, 20257 min read
Chain-of-Thought Prompting: Teaching AI to Show Its Work

Chain-of-Thought Prompting: Teaching AI to Show Its Work

A single phrase, 'let's think step by step', dramatically improved LLM reasoning. Explore the science behind chain-of-thought and where it's headed.

Oct 6, 20257 min read
The Multimodal Moment: When AI Learned to See and Speak

The Multimodal Moment: When AI Learned to See and Speak

Vision-language models can now describe images, read charts, and watch videos. Here's how the architecture enables true cross-modal understanding.

Sep 29, 20255 min read
Structured Outputs: Making LLMs Reliable in Production

Structured Outputs: Making LLMs Reliable in Production

Free-form text works fine in a chat window, but production systems need JSON. A walkthrough of constrained decoding, schema enforcement, and a worked Gemini example pairing structured output with function calling.

Sep 22, 20259 min read
AI Agents: Beyond Chat to Autonomous Action

AI Agents: Beyond Chat to Autonomous Action

LLMs were just the beginning. Explore how autonomous AI agents use tools, memory, and planning to act, not just respond.

Sep 15, 20257 min read

46 of 46 articles

About

I'm Oswaldo Orona, a Principal Database Administrator and AI practitioner based in Denver, CO. I hold an MCS-DS from UIUC (Tau Beta Pi, Phi Kappa Phi) and bring 25+ years of database experience alongside deep hands-on work in machine learning and AI engineering.

My focus areas include Retrieval-Augmented Generation (RAG), AI agents, Model Context Protocol (MCP), geospatial AI, and financial AI, all running in a self-hosted Proxmox home lab with Docker and LXC.

This blog is where I document explorations in latent space: the ideas, experiments, and systems that live between the data and the model.

@ooronaView repos

ML & AI

  • PyTorch
  • TensorFlow
  • RAG
  • NLP
  • Computer Vision
  • LLMs

Databases

  • PostgreSQL
  • Oracle
  • Redis
  • pgvector
  • DynamoDB

Infrastructure

  • Docker
  • Proxmox
  • Ansible
  • AWS
  • Linux

Languages

  • Python
  • R
  • SQL
  • PL/SQL
  • Java
  • Bash

Ask about Oswaldo

Explore my experience through a conversation. Ask how I built a project, find out about my background, or discuss where my skills fit your team.

You’re chatting with an AI assistant. It can help arrange a call or pass along a message. A conversation summary may be sent to Oswaldo.