latentSource

Mapping the Star Wars Universe

A fun data experiment: embedding Star Wars characters in 3D space based on their Force sensitivity and factions.

·5 min read
Share
Mapping the Star Wars Universe

Mapping the Star Wars universe

Data science doesn't always have to be serious. Sometimes the cleanest way to demonstrate a technique is to point it at something you actually care about. So I'm taking sentence embeddings, similarity graphs, and dimensionality reduction, and aiming them at a galaxy far, far away.

The question I'm asking: can we embed Star Wars characters into a vector space and find meaningful structure in their relationships? Spoiler — yes, and what falls out says something about how the franchise's narrative is wired.

Building character descriptions

The first step is writing text descriptions that capture each character's essential traits. I'm not going for exhaustive Wookieepedia entries. I want concise summaries that encode what we care about: faction, role, Force sensitivity, era, and key relationships.

characters = { "Luke Skywalker": "Jedi Knight, son of Anakin Skywalker, strong in the Force, " "hero of the Rebellion, trained by Obi-Wan and Yoda, redeemed his father", "Darth Vader": "Sith Lord, formerly Anakin Skywalker, fallen Jedi, serves the Emperor, " "father of Luke and Leia, mechanical suit, immense Force power", "Han Solo": "Smuggler, captain of the Millennium Falcon, no Force sensitivity, " "pragmatic, allied with the Rebellion, close friend of Chewbacca", "Leia Organa": "Princess of Alderaan, leader of the Rebellion, Force sensitive, " "twin sister of Luke, diplomat and military strategist", "Emperor Palpatine": "Sith Master, ruler of the Galactic Empire, manipulator, " "immense dark side power, orchestrated the Clone Wars", "Yoda": "Grand Master of the Jedi Order, 900 years old, immense Force wisdom, " "trained generations of Jedi, exiled on Dagobah", "Obi-Wan Kenobi": "Jedi Master, trained Anakin Skywalker, guardian of Luke, " "wise and disciplined, killed by Darth Vader on the Death Star", "Boba Fett": "Bounty hunter, clone of Jango Fett, no Force sensitivity, " "works for the highest bidder, Mandalorian armor", "Ahsoka Tano": "Former Jedi Padawan of Anakin Skywalker, left the Jedi Order, " "Force sensitive, independent warrior, fought in the Clone Wars", "Din Djarin": "Mandalorian bounty hunter, no Force sensitivity, foundling, " "protector of Grogu, follows the Mandalorian creed", "Kylo Ren": "Dark side Force user, son of Han Solo and Leia Organa, " "formerly Ben Solo, trained by Luke then Snoke, conflicted", "Rey": "Jedi, granddaughter of Palpatine, strong in the Force, scavenger from Jakku, " "trained by Luke and Leia, chose the Skywalker name", "Mace Windu": "Jedi Master, member of the Jedi Council, invented Vaapad lightsaber form, " "strong in the Force, confronted Palpatine", "Count Dooku": "Sith apprentice, formerly a Jedi Master, trained by Yoda, " "leader of the Separatists, elegant duelist", "Chewbacca": "Wookiee warrior, co-pilot of the Millennium Falcon, loyal companion " "of Han Solo, fought in the Clone Wars and Galactic Civil War", }

I wrote them deliberately to lean on relational and factional attributes. "Trained by Yoda" creates a link between two characters. "No Force sensitivity" is an important negative feature that shows up in the embedding.

Generating embeddings

A sentence embedding model converts each description into a dense vector. Sentence-BERT models (or successors like E5 or BGE) are right for this — cosine similarity in their output space tracks semantic similarity well.

from sentence_transformers import SentenceTransformer import numpy as np model = SentenceTransformer('all-MiniLM-L6-v2') names = list(characters.keys()) descriptions = list(characters.values()) embeddings = model.encode(descriptions) # Check similarity between two characters from sklearn.metrics.pairwise import cosine_similarity sim_matrix = cosine_similarity(embeddings)

The similarity matrix already tells a story. Luke's nearest neighbors will be Obi-Wan, Yoda, and Leia. Darth Vader sits next to Palpatine and Kylo Ren. Han and Chewbacca pair up, near Boba Fett and Din Djarin.

Building a similarity graph

A raw similarity matrix is hard to read. A graph makes the relationships visual and lets you actually explore them.

import networkx as nx G = nx.Graph() threshold = 0.45 # Only draw edges for meaningfully similar characters for i, name_i in enumerate(names): G.add_node(name_i) for j, name_j in enumerate(names): if i < j and sim_matrix[i][j] > threshold: G.add_edge(name_i, name_j, weight=float(sim_matrix[i][j]))

The threshold controls density. Too low and everything connects to everything. Too high and the graph falls apart into islands. I usually start at the median similarity and adjust until the picture is readable.

Visualizing in 2D and 3D

For the spatial view, drop the embedding dimensions with UMAP:

import umap import matplotlib.pyplot as plt reducer = umap.UMAP(n_components=2, n_neighbors=5, min_dist=0.3, metric='cosine', random_state=42) coords_2d = reducer.fit_transform(embeddings) # Color by faction faction_colors = { 'Jedi': '#4488ff', 'Sith': '#ff4444', 'Neutral': '#88cc88', 'Mandalorian': '#ffaa00' } factions = { 'Luke Skywalker': 'Jedi', 'Darth Vader': 'Sith', 'Han Solo': 'Neutral', 'Leia Organa': 'Jedi', 'Emperor Palpatine': 'Sith', 'Yoda': 'Jedi', 'Obi-Wan Kenobi': 'Jedi', 'Boba Fett': 'Mandalorian', 'Ahsoka Tano': 'Jedi', 'Din Djarin': 'Mandalorian', 'Kylo Ren': 'Sith', 'Rey': 'Jedi', 'Mace Windu': 'Jedi', 'Count Dooku': 'Sith', 'Chewbacca': 'Neutral', } fig, ax = plt.subplots(figsize=(12, 9)) for name, coord in zip(names, coords_2d): color = faction_colors[factions[name]] ax.scatter(coord[0], coord[1], c=color, s=120, edgecolors='white', zorder=2) ax.annotate(name, (coord[0], coord[1]), fontsize=8, ha='center', va='bottom') ax.set_title('Star Wars Characters in Embedding Space') plt.tight_layout() plt.savefig('star_wars_embeddings.png', dpi=150)

What the clusters reveal

A few patterns show up consistently in the 2D projection.

The Force axis. Force-sensitive characters (Jedi and Sith) sit in one region. Non-Force users (Han, Chewie, Boba Fett, Din Djarin) sit in another. It's the strongest axis of separation, which makes sense — Force sensitivity is the defining trait in the franchise and it's the most mentioned attribute in the descriptions.

Light vs dark within Force users. Jedi and Sith split cleanly. Yoda, Obi-Wan, and Mace Windu cluster tight. Palpatine and Dooku form their own pocket. The interesting characters are the ones in between. Ahsoka, who left the Jedi Order, sits slightly off the core Jedi cluster. Kylo Ren, described as "conflicted," drifts toward the boundary.

The gray zone. Han, Chewbacca, Boba Fett, and Din Djarin form a cluster of pragmatic non-Force users. Inside that cluster, the bounty hunters (Boba and Din) are closer to each other than to the Rebellion-aligned characters (Han and Chewie). Faction matters even among the "neutrals."

Cross-era bridges. Darth Vader often lands between the Jedi and Sith clusters because his description carries both identities. He's the narrative bridge of the saga, and the embedding picks that up. Count Dooku, described as "formerly a Jedi Master," pulls slightly toward the light side compared to pure Sith like Palpatine.

What this actually demonstrates

It's a toy example, but the pipeline is real. Embed text descriptions, compute similarities, visualize with UMAP, build a graph — that exact recipe applies to:

  • Product catalogs: embed product descriptions to find similar items and gaps in inventory
  • Research papers: map a field's literature to find clusters of related work and underexplored connections
  • Job postings: embed role descriptions to see how a company's hiring reflects its strategy
  • Customer support tickets: embed tickets to surface recurring issue categories without manual tagging

The Star Wars set is small enough to inspect by hand and rich enough to show real structure. Every cluster, every outlier, every bridge character matches a narrative truth about the franchise. That's what embeddings are good at: they recover structure you already knew was there, and sometimes structure you didn't.

May the Force be with your data.