JobMap
Describe the work you want to do, explore matching job postings, and inspect the evidence behind Jev's relevance judgments.
Explore the agent →I'm Oswaldo Orona. I write about the systems I build, from retrieval and AI agents to the data work underneath. Read the decisions, explore the results, and inspect the code.
Systems I built, with code and decisions you can inspect.
Describe the work you want to do, explore matching job postings, and inspect the evidence behind Jev's relevance judgments.
Explore the agent →An interactive explorer for fifteen character descriptions, with similarity scores, projection views, and a way to test what changes when you remove their named relationships.
Explore the agent →A movie and music search agent that turns a remembered line into a playable video, with the tool calls visible beside the result.
Explore the agent →A database transfer workflow that masks customer fields before rows reach an agent or its mailbox.
Explore the agent →A conversational transit map that resolves places and asks a routing engine to calculate the journey.
Explore the agent →An AI participant in a World Cup prediction pool, with saved forecasts and a leaderboard scored alongside human players.
Explore the agent →A geospatial investigation built to challenge the Mexican government’s account of the oil spill using vessel records, satellite layers, and dated statements.
Explore the agent →A public MCP endpoint that lets an assistant search the blog, browse its catalog, and retrieve complete articles.
Explore the agent →Projects, analysis, and explanations of the systems I work with.
I compared Jev with three local embedding models on 237 job postings. Jev led on graded relevance, but a stricter definition of a good match changed the conclusion.
A Gemini agent searches for movie scenes and songs, shows its tool calls, and returns a YouTube result inside the app.
A Gemini agent connects to MCP tools backed by PostGIS and pgRouting. The interesting part is resolving a place, choosing stations, and showing where the answer came from.
I combined vessel records, satellite layers, and a timeline in a geospatial app. Then I checked which parts of the animation were observations and which were interpolation.
I exposed the blog’s retrieval system as three read-only MCP tools. The implementation is small enough to inspect, and the same search tool appears in recorded agent sessions.
Ollama is the least painful way to serve local models on a single GPU, but the defaults are tuned for demos, not for work. What actually matters (quantization, context memory, keep-alive) and how to get a fine-tuned model of your own running behind the same API.
A grounded prediction agent competed under the same scoring rules as people. The published group-stage chart reports 47 correct outcomes from 72 picks; here is the implementation and what that result still needs to establish.
I put masking inside the producer’s data tool, then limited the tools that agent could call. Here is the code, the synthetic example, and the boundary this proof of concept actually provides.
A walkthrough of one real ReAct agent run. The query was "Houston, we have a problem" and the agent returned the exact YouTube timestamp where Tom Hanks delivers the line. You can replay the recorded run right inside this article, thinking tokens and all.
Everyone said prompt engineering was a fad. They were wrong. It just evolved. From artisanal prompting to systematic prompt design for production systems.
The architecture of this site: bilingual articles, recorded LLM sessions, hybrid retrieval, and a read-only MCP interface. Here is what the implementation actually does.
Guaranteeing valid output from LLMs requires more than prompting. Grammar-constrained decoding enforces structure at the token level. Here's how it works.
Cloud APIs are convenient but expensive. Explore how to run open-source LLMs on your own servers, from hardware selection to inference optimization.
You don't need a separate vector database. pgvector turns PostgreSQL into a semantic search engine, with HNSW indexes, hybrid queries, and full SQL power.
Claude Code brings an AI agent directly into your terminal. Explore what autonomous coding tools mean for software engineering workflows.
Single-prompt AI is hitting its ceiling. Agentic workflows chain multiple LLM calls with tools, branching, and feedback loops to tackle complex tasks reliably.
The Model Context Protocol standardizes how AI models discover and use tools. Here's how MCP servers work and why they matter for the agentic future.
Large models are powerful but expensive. Distillation transfers their knowledge into smaller, faster, cheaper models, and the results are surprisingly good.
Gemini 1.5 Pro pushed context to 1M tokens. Claude 3.5 followed. But is a bigger context window always better, and what does it really enable?
Hallucinations happen when models guess. Knowledge graphs give LLMs a structured, verifiable backbone, and the combination is more powerful than either alone.
From Apple's Neural Engine to Qualcomm's NPUs, AI is moving off the cloud and onto your device. Here's what's driving the shift and what it unlocks.
Both approaches customize LLMs for specific domains, but they solve different problems. Here's how to pick the right tool, and when to combine them.
Inference is slow because tokens are generated one at a time. Speculative decoding breaks that constraint, without changing the model's output distribution.
Copilot, Cursor, Claude Code, Codex. AI assistants write code faster than ever. But are developers getting better or just more dependent?
Extended thinking, test-time compute, and chain-of-thought at training time, unpacking how a new class of models trades latency for accuracy.
A Chinese lab released a reasoning model that matched OpenAI o1 at a fraction of the cost, and then open-sourced it. Here's why that matters enormously.
From reasoning-native models to autonomous software engineers, here are the trends that will define AI in 2026.
A look back at the breakthroughs, surprises, and pivots that defined artificial intelligence in 2025, from reasoning models to the agent revolution.
Latent spaces encode meaning in geometry. Learn how dimensionality reduction techniques like t-SNE and UMAP make these invisible structures visible.
A practical guide to building a Retrieval-Augmented Generation pipeline using Google's Gemini API, File Search Store, and embedding models.
Optimize your ML infrastructure with strategies like MIG and quantization to cut costs without sacrificing performance.
A step-by-step technical guide to constructing a Retrieval-Augmented Generation system for enterprise data.
A visual breakdown of how AI is surging across industries, from healthcare to finance, over the next five years.
Track the exponential drop in token generation costs and what this democratization means for the future of software.
Understand the power laws governing AI performance: why bigger models and more compute consistently yield better results.
Trace the lineage of neural networks from the humble Perceptron to the massive ResNets feeding today's AI.
Peel back the layers of the Transformer architecture to visualize how attention heads process information.
A fun data experiment: embedding fifteen Star Wars characters, testing which clusters actually show up, and finding out the model sorts them by rank and story arc instead of light and dark.
Learn how K-Means clustering uncovers hidden patterns in user behavior to drive targeted business strategies.
Explore the math behind language: see how 'King - Man + Woman = Queen' plays out visually in vector space.
Unlock the secrets of high-dimensional data compression and how mapping complex inputs to simpler spaces drives generative AI.
Embeddings are the language of semantic search, and vector databases are where they live. Learn how ANN indices power retrieval at scale.
A single phrase, 'let's think step by step', dramatically improved LLM reasoning. Explore the science behind chain-of-thought and where it's headed.
Vision-language models can now describe images, read charts, and watch videos. Here's how the architecture enables true cross-modal understanding.
Free-form text works fine in a chat window, but production systems need JSON. A walkthrough of constrained decoding, schema enforcement, and a worked Gemini example pairing structured output with function calling.
LLMs were just the beginning. Explore how autonomous AI agents use tools, memory, and planning to act, not just respond.
46 of 46 articles
I'm Oswaldo Orona, a Principal Database Administrator and AI practitioner based in Denver, CO. I hold an MCS-DS from UIUC (Tau Beta Pi, Phi Kappa Phi) and bring 25+ years of database experience alongside deep hands-on work in machine learning and AI engineering.
My focus areas include Retrieval-Augmented Generation (RAG), AI agents, Model Context Protocol (MCP), geospatial AI, and financial AI, all running in a self-hosted Proxmox home lab with Docker and LXC.
This blog is where I document explorations in latent space: the ideas, experiments, and systems that live between the data and the model.
Explore my experience through a conversation. Ask how I built a project, find out about my background, or discuss where my skills fit your team.
You’re chatting with an AI assistant. It can help arrange a call or pass along a message. A conversation summary may be sent to Oswaldo.