<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Rushi Patel]]></title><description><![CDATA[Rushi Patel]]></description><link>https://rsp180002.hashnode.dev</link><image><url>https://cdn.hashnode.com/res/hashnode/image/upload/v1593680282896/kNC7E8IR4.png</url><title>Rushi Patel</title><link>https://rsp180002.hashnode.dev</link></image><generator>RSS for Node</generator><lastBuildDate>Wed, 07 Oct 2026 14:32:22 GMT</lastBuildDate><atom:link href="https://rsp180002.hashnode.dev/rss.xml" rel="self" type="application/rss+xml"/><language><![CDATA[en]]></language><ttl>60</ttl><item><title><![CDATA[Building an Explainable AI Learning Path Recommender with Embeddings, Knowledge Graphs and LLMs]]></title><description><![CDATA[From simple keyword matching to a testable, prerequisite-aware recommendation pipeline, with a local open-source LLM that explains every pick.
Tutorial for Python developers. Uses free, local tools on]]></description><link>https://rsp180002.hashnode.dev/building-an-explainable-ai-learning-path-recommender-with-embeddings-knowledge-graphs-and-llms</link><guid isPermaLink="true">https://rsp180002.hashnode.dev/building-an-explainable-ai-learning-path-recommender-with-embeddings-knowledge-graphs-and-llms</guid><dc:creator><![CDATA[Rushi]]></dc:creator><pubDate>Mon, 05 Oct 2026 05:36:47 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6ababed9a0f54543c7636b38/106fb069-3c1e-4e76-8ac5-2e6613a84bbf.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>From simple keyword matching to a testable, prerequisite-aware recommendation pipeline, with a local open-source LLM that explains every pick.</p>
<p>Tutorial for Python developers. Uses free, local tools only. About 15 minutes to read.</p>
<p>Netflix recommends movies, Spotify recommends songs, and Amazon recommends products. But recommendation systems can also answer a much more personal question:</p>
<blockquote>
<p><strong>What should I learn next to reach my career goal?</strong></p>
</blockquote>
<p>A developer who knows Python and SQL and wants to become a Generative AI Engineer doesn't need a random list of AI courses. A useful system has to understand what they already know, what the target role requires, which skills are missing, which skills depend on others, and in what order to learn them.</p>
<p>It's tempting to just ask an LLM <em>"What should I learn to become an AI Engineer?"</em> The answer will sound confident, but it can't be tested, it changes every time, and nothing guarantees it respects prerequisites. So in this project I built a real recommendation pipeline instead, and used the LLM only where it shines: understanding the question and explaining the results.</p>
<p>The stack: <strong>Python, pandas, Sentence Transformers, NetworkX, multi-factor ranking, ranking metrics, and a local open-source LLM with tool calling</strong> through Ollama, so the whole project runs free on your own machine.</p>
<h2>The architecture</h2>
<p>User profile Skill normalization Skill gap analysis Embeddings: relevance Knowledge graph: order Multi-factor ranking LLM: understand and explain deterministic engine decides, the LLM only explains</p>
<p>The recommendation pipeline. Everything above the coral box is testable code.</p>
<p>The most important design decision: <strong>the LLM does not control the recommendations.</strong> The deterministic engine finds gaps, orders skills, and ranks resources. The LLM receives structured results and turns them into a clear explanation.</p>
<p>Install everything first:</p>
<pre><code class="language-bash">pip install sentence-transformers networkx numpy ollama
</code></pre>
<h2>Step 1: A small career dataset</h2>
<p>For this prototype I wrote a role-to-skill dataset by hand. In production you'd replace it with skills extracted from real job postings, or with an occupational taxonomy such as <strong>O*NET</strong> (US Department of Labor) or <strong>ESCO</strong> (European Commission), both free to use.</p>
<pre><code class="language-python">roles = {
    "Generative AI Engineer": [
        "Python", "Machine Learning", "Deep Learning", "PyTorch", "Transformers",
        "Hugging Face", "Embeddings", "Vector Databases", "RAG", "LangChain",
        "LangGraph", "FastAPI", "LLM Evaluation",
    ],
    "Machine Learning Engineer": [
        "Python", "Statistics", "Pandas", "NumPy", "Scikit-learn", "Machine Learning",
        "Deep Learning", "PyTorch", "MLOps",
    ],
    "Data Scientist": [
        "Python", "SQL", "Statistics", "Pandas", "NumPy", "Matplotlib",
        "Scikit-learn", "Machine Learning",
    ],
}
ALL_SKILLS = sorted({s for skills in roles.values() for s in skills})
</code></pre>
<p>Our example user:</p>
<pre><code class="language-python">profile = {
    "current_skills": ["Python", "JavaScript", "SQL", "React"],
    "target_role": "Generative AI Engineer",
    "experience_level": "intermediate",
    "hours_per_week": 10,
}
</code></pre>
<h2>Step 2: Find the skill gaps</h2>
<p>The baseline is simple set logic: what the role requires, minus what the user knows.</p>
<pre><code class="language-python">def find_skill_gaps(current_skills, target_role):
    known = {s.lower().strip() for s in current_skills}
    return [s for s in roles.get(target_role, []) if s.lower() not in known]

print(find_skill_gaps(profile["current_skills"], profile["target_role"]))
</code></pre>
<pre><code class="language-text">['Machine Learning', 'Deep Learning', 'PyTorch', 'Transformers', 'Hugging Face',
 'Embeddings', 'Vector Databases', 'RAG', 'LangChain', 'LangGraph', 'FastAPI',
 'LLM Evaluation']
</code></pre>
<p>This works, but only if the user types skill names <em>exactly</em> the way our dataset does.</p>
<h2>Step 3: Understand messy skill names with embeddings</h2>
<p>Real users write "ML", "stats", "NLP", or "building apps with LLMs". Exact matching treats all of these as unknown. <strong>Embeddings</strong> fix most of this by turning text into vectors where similar meanings sit close together, so we can map free-text skills to our canonical list.</p>
<p>Embeddings are weak at short acronyms, though, so a good system combines them with a small alias table. That's a common real-world pattern: rules for the cases you know, embeddings for everything else.</p>
<pre><code class="language-python">import numpy as np
from sentence_transformers import SentenceTransformer

embedding_model = SentenceTransformer("all-MiniLM-L6-v2")

ALIASES = {"ml": "Machine Learning", "dl": "Deep Learning", "stats": "Statistics",
           "hf": "Hugging Face", "vector db": "Vector Databases",
           "retrieval-augmented generation": "RAG", "llm evals": "LLM Evaluation"}

def normalize_skills(raw_skills, model, threshold=0.6):
    canon_emb = model.encode(ALL_SKILLS, normalize_embeddings=True)
    result = {}
    for raw in raw_skills:
        key = raw.lower().strip()
        exact = next((s for s in ALL_SKILLS if s.lower() == key), None)
        if exact or key in ALIASES:
            result[raw] = exact or ALIASES[key]
            continue
        query = model.encode([raw], normalize_embeddings=True)[0]
        scores = canon_emb @ query            # cosine similarity (vectors are normalized)
        best = int(np.argmax(scores))
        result[raw] = ALL_SKILLS[best] if scores[best] &gt;= threshold else None
    return result

print(normalize_skills(["ml", "stats", "deep neural networks", "cooking"], embedding_model))
</code></pre>
<p>The <code>threshold</code> matters: without it, "cooking" would be forced onto whichever skill happens to be closest. With it, unrelated text returns <code>None</code>, and the app can simply ask the user what they meant. Tune the threshold on a few examples of your own.</p>
<h2>Step 4: Model prerequisites with a knowledge graph</h2>
<p>Similarity can't tell you that you need Machine Learning before PyTorch. A semantic recommender might happily suggest LangGraph first, because it's strongly related to "Generative AI Engineer".</p>
<p>So I modeled prerequisites as a <strong>directed graph</strong>, where an edge <code>A → B</code> means "learn A before B". Unlike a simple chain, a real skill graph branches and merges: RAG needs both Embeddings and Vector Databases, and Machine Learning needs both Python and Statistics.</p>
<pre><code class="language-python">import networkx as nx

G = nx.DiGraph()
G.add_edges_from([
    ("Python", "Machine Learning"), ("Statistics", "Machine Learning"),
    ("Python", "Pandas"), ("Python", "NumPy"), ("Python", "FastAPI"),
    ("Machine Learning", "Deep Learning"), ("Deep Learning", "PyTorch"),
    ("PyTorch", "Transformers"), ("Transformers", "Hugging Face"),
    ("Transformers", "Embeddings"), ("Embeddings", "Vector Databases"),
    ("Embeddings", "RAG"), ("Vector Databases", "RAG"),
    ("RAG", "LangChain"), ("LangChain", "LangGraph"), ("RAG", "LLM Evaluation"),
])

assert nx.is_directed_acyclic_graph(G), "Prerequisites must not form a loop"
</code></pre>
<p>The <code>assert</code> is a small but important safety check. If someone accidentally adds a loop (A needs B, B needs A), no valid learning order exists.</p>
<h2>Step 5: Build a personalized learning path</h2>
<p>Here's where the graph earns its place. For every missing skill, we also collect its <strong>ancestors</strong>, meaning all the prerequisites that lead to it. Then we sort everything with a <strong>topological sort</strong>, which guarantees every skill appears after its prerequisites.</p>
<pre><code class="language-python">def learning_path(current_skills, target_role):
    known = {s.lower() for s in current_skills}
    needed = set()
    for skill in find_skill_gaps(current_skills, target_role):
        needed.add(skill)
        if skill in G:
            needed |= nx.ancestors(G, skill)      # pull in hidden prerequisites
    needed = {s for s in needed if s.lower() not in known}
    order = list(nx.lexicographical_topological_sort(G))
    in_graph = [s for s in order if s in needed]
    return in_graph + sorted(needed - set(in_graph))

print(learning_path(profile["current_skills"], profile["target_role"]))
</code></pre>
<p>Two things in this output are worth noticing:</p>
<ul>
<li><p><strong>Statistics appears, even though the role never listed it.</strong> The graph discovered it as a hidden prerequisite of Machine Learning. Simple gap analysis would have missed it.</p>
</li>
<li><p><strong>FastAPI comes first.</strong> It only depends on Python, which the user already knows, so it's a quick win they can start today. The sort is deterministic, so the same profile always gets the same path.</p>
</li>
</ul>
<h2>Step 6: Measure readiness</h2>
<p>For each skill, <strong>readiness</strong> is the share of its prerequisites the user already knows. A score of 1 means "you can start now", and 0 means "not yet".</p>
<pre><code class="language-python">def readiness(skill, known_skills):
    if skill not in G:
        return 1.0
    prereqs = list(G.predecessors(skill))
    if not prereqs:
        return 1.0
    known = {s.lower() for s in known_skills}
    return sum(p.lower() in known for p in prereqs) / len(prereqs)

print(readiness("Machine Learning", profile["current_skills"]))   # 0.5: has Python, missing Statistics
print(readiness("RAG", profile["current_skills"])) 
</code></pre>
<h2>Step 7: A learning resource catalog</h2>
<p>Now we need things to recommend. Each resource covers some skills and has a difficulty level and an estimated number of hours.</p>
<pre><code class="language-python">resources = [
    {"id": 1, "title": "Statistics for Machine Learning", "skills": ["Statistics"], "difficulty": "beginner", "hours": 15},
    {"id": 2, "title": "Machine Learning Fundamentals", "skills": ["Machine Learning", "Scikit-learn"], "difficulty": "beginner", "hours": 30},
    {"id": 3, "title": "Deep Learning with PyTorch", "skills": ["Deep Learning", "PyTorch"], "difficulty": "intermediate", "hours": 40},
    {"id": 4, "title": "Transformers and Hugging Face", "skills": ["Transformers", "Hugging Face"], "difficulty": "intermediate", "hours": 25},
    {"id": 5, "title": "Embeddings and Vector Search", "skills": ["Embeddings", "Vector Databases"], "difficulty": "intermediate", "hours": 12},
    {"id": 6, "title": "Building RAG Applications", "skills": ["RAG", "Vector Databases"], "difficulty": "intermediate", "hours": 20},
    {"id": 7, "title": "Agentic AI with LangGraph", "skills": ["LangGraph", "LangChain"], "difficulty": "advanced", "hours": 20},
    {"id": 8, "title": "Shipping ML APIs with FastAPI", "skills": ["FastAPI"], "difficulty": "beginner", "hours": 8},
    {"id": 9, "title": "Evaluating LLM Applications", "skills": ["LLM Evaluation", "RAG"], "difficulty": "advanced", "hours": 10},
]
</code></pre>
<h2>Step 8: Rank with multiple signals</h2>
<p>Real recommenders rarely rely on one score. I combine five signals, each between 0 and 1:</p>
<pre><code class="language-text">Score = 0.35 × semantic relevance    (how related is it to the goal?)
      + 0.25 × gap coverage          (how much of it is new to me?)
      + 0.20 × readiness             (do I have the prerequisites?)
      + 0.10 × difficulty match      (is it at my level?)
      + 0.10 × time fit              (can I finish it in about 4 weeks?)
</code></pre>
<pre><code class="language-python">LEVELS = {"beginner": 0, "intermediate": 1, "advanced": 2}
WEIGHTS = {"semantic": 0.35, "gap": 0.25, "readiness": 0.20,
           "difficulty": 0.10, "time_fit": 0.10}

def signals(resource, profile, missing):
    skills = resource["skills"]
    gap = sum(s in missing for s in skills) / len(skills)
    ready = float(np.mean([readiness(s, profile["current_skills"]) for s in skills]))
    diff = 1 - abs(LEVELS[resource["difficulty"]] - LEVELS[profile["experience_level"]]) / 2
    weeks = resource["hours"] / profile["hours_per_week"]
    time_fit = 1.0 if weeks &lt;= 4 else 4 / weeks
    return {"gap": gap, "readiness": ready, "difficulty": diff, "time_fit": time_fit}

def rank_resources(profile, model, top_k=5):
    missing = set(learning_path(profile["current_skills"], profile["target_role"]))
    goal = f"{profile['target_role']}: " + ", ".join(sorted(missing))
    texts = [r["title"] + ". Skills: " + ", ".join(r["skills"]) for r in resources]
    res_emb = model.encode(texts, normalize_embeddings=True)   # encode all at once, not in a loop
    goal_emb = model.encode([goal], normalize_embeddings=True)[0]
    semantic = (res_emb @ goal_emb + 1) / 2                    # rescale cosine to [0, 1]
    ranked = []
    for r, sem in zip(resources, semantic):
        if not set(r["skills"]) &amp; missing:
            continue                                           # nothing new to learn here
        reasons = {"semantic": float(sem), **signals(r, profile, missing)}
        score = sum(WEIGHTS[k] * v for k, v in reasons.items())
        ranked.append({"title": r["title"], "score": round(score, 3),
                       "reasons": {k: round(v, 2) for k, v in reasons.items()}})
    return sorted(ranked, key=lambda x: x["score"], reverse=True)[:top_k]

for item in rank_resources(profile, embedding_model):
    print(item)
</code></pre>
<p>Here are the four non-semantic signals for our example user, straight from the code. Your final scores will also include the semantic signal from the embedding model.</p>
<pre><code class="language-text">resource                         gap   ready  diff  time
Statistics for Machine Learning  1.00  1.00   0.50  1.00
Machine Learning Fundamentals    0.50  0.75   0.50  1.00
Deep Learning with PyTorch       1.00  0.00   1.00  1.00
Transformers and Hugging Face    1.00  0.00   1.00  1.00
Embeddings and Vector Search     1.00  0.00   1.00  1.00
Building RAG Applications        1.00  0.00   1.00  1.00
Agentic AI with LangGraph        1.00  0.00   0.50  1.00
Shipping ML APIs with FastAPI    1.00  1.00   0.50  1.00
Evaluating LLM Applications      1.00  0.00   0.50  1.00
</code></pre>
<p>This table explains the system's behavior. "Agentic AI with LangGraph" is highly relevant to the goal, but its readiness is 0, so it can't jump to the top. "Statistics" and "FastAPI" score full readiness, so they surface as sensible first steps. Every number is also stored in <code>reasons</code>, which makes each recommendation <strong>explainable</strong> by design.</p>
<p>Note that the weights are a starting point, not a truth. Once you collect real feedback, you can learn them from data instead of choosing them by hand.</p>
<h2>Step 9: Add a local LLM with tool calling</h2>
<p>At this point the system works without any LLM, and that's intentional. Now we add one for what it does best: understanding a free-form question and explaining structured results.</p>
<p>I use <strong>Ollama</strong>, which runs open-source models like Llama 3.1 locally, for free, with no API key, and no user data leaving the machine. Install Ollama from <a href="https://ollama.com">ollama.com</a>, then run <code>ollama pull llama3.1</code>.</p>
<p>We expose the whole engine as a single <strong>tool</strong>. The model decides when to call it and fills in the arguments, but the recommendations themselves always come from our code.</p>
<pre><code class="language-python">import json
import ollama

def recommend_learning_path(current_skills, target_role,
                            experience_level="intermediate", hours_per_week=10):
    profile = {"current_skills": current_skills, "target_role": target_role,
               "experience_level": experience_level, "hours_per_week": hours_per_week}
    return {
        "skill_gaps": find_skill_gaps(current_skills, target_role),
        "learning_path": learning_path(current_skills, target_role),
        "resources": rank_resources(profile, embedding_model),
    }

tools = [{
    "type": "function",
    "function": {
        "name": "recommend_learning_path",
        "description": "Return skill gaps, an ordered learning path and ranked "
                       "learning resources for a target role.",
        "parameters": {
            "type": "object",
            "properties": {
                "current_skills": {"type": "array", "items": {"type": "string"}},
                "target_role": {"type": "string", "enum": list(roles)},
                "experience_level": {"type": "string",
                                     "enum": ["beginner", "intermediate", "advanced"]},
                "hours_per_week": {"type": "number"},
            },
            "required": ["current_skills", "target_role"],
        },
    },
}]

SYSTEM = (
    "You are a career learning assistant. For any question about what to learn, "
    "call recommend_learning_path. Only recommend skills and resources the tool "
    "returns, keep the tool's order, and explain each pick using its 'reasons' "
    "scores. If a resource has low readiness, say which prerequisite comes first."
)

def ask(question, model="llama3.1"):
    messages = [{"role": "system", "content": SYSTEM},
                {"role": "user", "content": question}]
    response = ollama.chat(model=model, messages=messages, tools=tools)
    messages.append(response.message)
    for call in response.message.tool_calls or []:
        if call.function.name == "recommend_learning_path":
            result = recommend_learning_path(**call.function.arguments)
            messages.append({"role": "tool", "content": json.dumps(result)})
    final = ollama.chat(model=model, messages=messages)
    return final.message.content

print(ask("I'm a React developer who knows Python and SQL. I can study about "
          "10 hours a week. How do I become a Generative AI Engineer?"))
</code></pre>
<p>Two guardrails are doing quiet but important work here. The <code>enum</code> on <code>target_role</code> stops the model from inventing a role we have no data for, and the system prompt tells it to explain only what the tool returned. Because the model runs locally, you can also swap in any other tool-capable model by changing one string.</p>
<h2>Step 10: Evaluate it properly</h2>
<p>A recommender shouldn't be judged by whether its output "looks good". Standard ranking metrics make quality measurable:</p>
<ul>
<li><p><strong>Precision@K:</strong> of the top K recommendations, how many were relevant?</p>
</li>
<li><p><strong>Recall@K:</strong> of all relevant items, how many made it into the top K?</p>
</li>
<li><p><strong>NDCG@K:</strong> rewards placing relevant items near the top, not just including them.</p>
</li>
</ul>
<pre><code class="language-python">def precision_at_k(recommended, relevant, k):
    return len(set(recommended[:k]) &amp; set(relevant)) / k

def recall_at_k(recommended, relevant, k):
    return len(set(recommended[:k]) &amp; set(relevant)) / len(relevant)

def ndcg_at_k(recommended, relevant, k):
    dcg = sum(1 / np.log2(i + 2) for i, item in enumerate(recommended[:k]) if item in relevant)
    ideal = sum(1 / np.log2(i + 2) for i in range(min(len(relevant), k)))
    return dcg / ideal

recommended = ["Machine Learning Fundamentals", "Agentic AI with LangGraph",
               "Statistics for Machine Learning", "Shipping ML APIs with FastAPI",
               "Deep Learning with PyTorch"]
relevant = ["Statistics for Machine Learning", "Machine Learning Fundamentals",
            "Deep Learning with PyTorch"]   # e.g. what a human mentor would pick first

print(precision_at_k(recommended, relevant, 5))        # 0.6
print(round(recall_at_k(recommended, relevant, 3), 2)) # 0.67
print(round(ndcg_at_k(recommended, relevant, 5), 3))   # 0.885
</code></pre>
<p>Where do "relevant" labels come from? A practical option is to ask a few experienced engineers or mentors to build learning paths for sample profiles, then compare the system's output with theirs.</p>
<h3>An ablation study</h3>
<p>The experiment I find most interesting is comparing versions of the system, adding one component at a time:</p>
<table>
<thead>
<tr>
<th>Version</th>
<th>Approach</th>
</tr>
</thead>
<tbody><tr>
<td>Baseline</td>
<td>Exact keyword matching</td>
</tr>
<tr>
<td>A</td>
<td>Embedding similarity only</td>
</tr>
<tr>
<td>B</td>
<td>A + knowledge graph prerequisites</td>
</tr>
<tr>
<td>C</td>
<td>B + multi-factor ranking</td>
</tr>
<tr>
<td>D</td>
<td>C + LLM explanations (measured with user ratings of clarity and trust)</td>
</tr>
</tbody></table>
<p>The research question: <em>do prerequisite relationships and user-context signals improve recommendation quality compared with semantic similarity alone?</em> Note that version D shouldn't change the ranking at all, since the LLM only explains. That's exactly why it needs a different measure: whether people understand and trust the recommendations more.</p>
<h2>Beyond the prototype</h2>
<p>A few problems appear quickly once real people use a system like this:</p>
<ul>
<li><p><strong>Cold start:</strong> new users have no history, so the first recommendations must come from their profile, which is why this design is content-based. Collaborative signals ("people with your background found this course useful") can be added as interaction data grows.</p>
</li>
<li><p><strong>Diversity:</strong> ten nearly identical RAG tutorials aren't helpful. A diversity term can balance fundamentals, frameworks, projects, deployment, and evaluation.</p>
</li>
<li><p><strong>Feedback loop:</strong> when a user completes a resource or rates its difficulty, update their profile and re-rank. Their "known skills" grow, readiness scores rise, and the path naturally moves forward.</p>
</li>
<li><p><strong>Fairness:</strong> if you later learn from user behavior, check that recommendations don't steer groups of users toward lower-paying paths. Explainable scores make this much easier to audit.</p>
</li>
</ul>
<h2>Suggested project structure</h2>
<pre><code class="language-text">ai-career-recommender/
├── README.md
├── requirements.txt
├── data/          roles.json, skills_graph.json, resources.json
├── src/           skills.py, graph.py, ranking.py, recommender.py, evaluation.py
├── llm/           tools.py, assistant.py
├── notebooks/     experiments.ipynb
├── tests/         test_graph.py, test_ranking.py
└── app.py
</code></pre>
<p>Good first tests: the graph has no cycles, no skill ever appears before its prerequisites in a path, and a user never gets a recommendation for something they already know.</p>
<h2>What I learned</h2>
<p>The biggest lesson: adding an LLM doesn't automatically make a recommender intelligent. Each component solves a different problem. <strong>Embeddings</strong> handle meaning and messy input. The <strong>knowledge graph</strong> handles prerequisites and order. <strong>Ranking</strong> decides what comes first. <strong>Profiles</strong> provide personalization. And the <strong>LLM</strong> handles conversation and explanation. Combined, they make a system that's testable, explainable, and far more trustworthy than asking a chatbot for a list.</p>
<h2>Future improvements</h2>
<p>Natural next steps include extracting skills automatically from resumes, building the role dataset from real job postings or O*NET, moving the graph to Neo4j and resources to a vector database, learning ranking weights from feedback, adding portfolio projects and certifications, generating weekly study plans, and serving it all through FastAPI with a React frontend.</p>
<h2>Conclusion</h2>
<p>This project grew step by step: from keyword matching, to embeddings, to a prerequisite graph, to multi-factor ranking, to an LLM that explains it all. The principle that held it together was <strong>separating recommendation from generation</strong>. The engine decides what to recommend using data, graphs, and measurable rules. The LLM sits on top to understand people and explain the results.</p>
<p>In the next post, I plan to extend this into an agentic version, where specialized agents analyze the profile, skill gaps, resources, and job market before collaborating on a roadmap, and to test whether that actually beats the simpler pipeline.</p>
<p><em>This post is part of my series on practical, explainable recommendation systems. You may also like</em> <a href="https://rsp180002.hashnode.dev/building-an-ai-recommender-for-national-parks"><em>Building an AI Recommender for National Parks</em></a><em>.</em></p>
]]></content:encoded></item><item><title><![CDATA[Helping Women with PCOS Find Research Studies with AI]]></title><description><![CDATA[Building a grounded, privacy-aware study finder with ClinicalTrials.gov, embeddings and an LLM, with working Python code.
Tutorial for Python developers interested in AI for healthcare. About 12 minut]]></description><link>https://rsp180002.hashnode.dev/helping-women-with-pcos-find-research-studies-with-ai</link><guid isPermaLink="true">https://rsp180002.hashnode.dev/helping-women-with-pcos-find-research-studies-with-ai</guid><dc:creator><![CDATA[Rushi]]></dc:creator><pubDate>Sat, 03 Oct 2026 06:21:47 GMT</pubDate><content:encoded><![CDATA[<p>Building a grounded, privacy-aware study finder with ClinicalTrials.gov, embeddings and an LLM, with working Python code.</p>
<p>Tutorial for Python developers interested in AI for healthcare. About 12 minutes to read.</p>
<blockquote>
<p><strong>Please note:</strong> This article is about building software. The tool described here helps people <em>discover</em> research studies. It does not diagnose, treat, or give medical advice, and it cannot decide whether anyone qualifies for a study. Always talk with a doctor before joining a clinical trial.</p>
</blockquote>
<p>Polycystic ovary syndrome (PCOS) is one of the most common hormonal conditions. The World Health Organization estimates it affects 8 to 13 percent of women of reproductive age, and that up to 70 percent of cases are undiagnosed. Yet PCOS research still has many open questions, and one of the quiet reasons is that studies struggle to find participants.</p>
<p>On the other side, people living with PCOS often never hear about studies they could join. The information is public, but it's written for researchers: long eligibility lists, medical terms, and hundreds of studies to sort through.</p>
<p>That gap is a perfect job for a recommender. In this article we'll build a tool that takes a person's own description of their situation and finds PCOS studies that may be worth asking their doctor about. We'll use real, live data from <strong>ClinicalTrials.gov</strong>, the US National Library of Medicine's public registry.</p>
<p>If you read my earlier post on <a href="https://rsp180002.hashnode.dev/building-an-ai-recommender-for-national-parks">building an AI recommender for national parks</a>, you'll recognize the pattern: hard rules first, then semantic matching, and finally a language model that explains results without inventing anything.</p>
<h2>The plan</h2>
<ol>
<li><p>Download recruiting PCOS studies from the ClinicalTrials.gov API.</p>
</li>
<li><p>Flatten the nested data into a simple table.</p>
</li>
<li><p>Apply hard filters: age, sex, and location.</p>
</li>
<li><p>Rank the remaining studies by meaning using AI embeddings.</p>
</li>
<li><p>Use an LLM to explain each study in plain language, grounded only in its official text.</p>
</li>
<li><p>Wrap it in a small web app.</p>
</li>
<li><p>Design it responsibly.</p>
</li>
</ol>
<p>Install the libraries:</p>
<pre><code class="language-bash">pip install requests pandas sentence-transformers anthropic streamlit
</code></pre>
<h2>Step 1: Download PCOS studies</h2>
<p>The ClinicalTrials.gov API v2 is free and needs no API key. We ask for studies about PCOS that are recruiting now or will be soon. Results come in pages, so we follow the <code>nextPageToken</code> until there are no more.</p>
<pre><code class="language-python">import re
import requests
import pandas as pd

API_URL = "https://clinicaltrials.gov/api/v2/studies"

def fetch_pcos_trials(max_pages=5): 
    """Download recruiting PCOS studies from ClinicalTrials.gov.""" \
    params = { "query.cond": "polycystic ovary syndrome",              "filter.overallStatus": "RECRUITING|NOT_YET_RECRUITING", "pageSize": 100, } 
    studies = []
    for _ in range(max_pages):
        response =     requests.get(API_URL, params=params, timeout=30)
        response.raise_for_status()
        data = response.json()
        studies.extend(data.get("studies", []))
        token = data.get("nextPageToken")
        if not token:
            break
        params["pageToken"] = token
    return studies

studies = fetch_pcos_trials()
print(f"Downloaded {len(studies)} studies")
</code></pre>
<blockquote>
<p><strong>Be a good API citizen:</strong> the API is shared public infrastructure. Keep requests modest and save results locally instead of downloading on every run.</p>
</blockquote>
<h2>Step 2: Flatten the data</h2>
<p>Each study comes back as deeply nested JSON, split into modules like <code>identificationModule</code>, <code>eligibilityModule</code>, and <code>contactsLocationsModule</code>. We pull out only what we need. Ages arrive as text like <code>"18 Years"</code>, so a small helper converts them into numbers.</p>
<pre><code class="language-python">def parse_age(text):
    """Turn '18 Years' or '6 Months' into years. Missing means no limit."""
    if not text:
        return None
    number, unit = re.match(r"(\d+)\s*(\w+)", text).groups()
    number = float(number)
    unit = unit.lower()
    if unit.startswith("month"):
        return number / 12
    if unit.startswith("week"):
        return number / 52
    if unit.startswith("day"):
        return number / 365
    return number

def to_table(studies):
    """Flatten the nested API response into one row per study."""

    rows = []

    for study in studies:
        p = study.get("protocolSection", {})

        ident = p.get("identificationModule", {})
        desc = p.get("descriptionModule", {})
        elig = p.get("eligibilityModule", {})
        locs = p.get("contactsLocationsModule", {}).get("locations", [])

        rows.append({
            "nct_id": ident.get("nctId"),
            "title": ident.get("briefTitle", ""),
            "status": p.get("statusModule", {}).get("overallStatus"),
            "summary": desc.get("briefSummary", ""),
            "conditions": ", ".join(
                p.get("conditionsModule", {}).get("conditions", [])
            ),
            "min_age": parse_age(elig.get("minimumAge")),
            "max_age": parse_age(elig.get("maximumAge")),
            "sex": elig.get("sex", "ALL"),
            "criteria": elig.get("eligibilityCriteria", ""),
            "states": sorted({
                l.get("state")
                for l in locs
                if l.get("country") == "United States"
                and l.get("state")
            }),
            "url": f"[https://clinicaltrials.gov/study/{ident.get(](https://clinicaltrials.gov/study/{ident.get\()'nctId')}",
        })

    return pd.DataFrame(rows)


trials = to_table(studies)
</code></pre>
<h2>Step 3: Apply hard filters first</h2>
<p>Some rules are not a matter of opinion. If a study only accepts people aged 12 to 17, it should never be shown to a 28-year-old, no matter how well the description matches. So we remove those studies <strong>before</strong> any AI ranking happens.</p>
<pre><code class="language-python">def basic_filter(trials, age, state=None):
    """Keep studies whose age range fits, that accept women,
    and (optionally) have a site in the user's state."""
    ok_age = (trials["min_age"].isna() | (trials["min_age"] &lt;= age)) &amp; \
             (trials["max_age"].isna() | (trials["max_age"] &gt;= age))
    ok_sex = trials["sex"].isin(["FEMALE", "ALL"])
    result = trials[ok_age &amp; ok_sex]
    if state:
        result = result[result["states"].apply(lambda s: state in s)]
    return result.reset_index(drop=True)

candidates = basic_filter(trials, age=28, state="Texas")
</code></pre>
<p>Here's how the filter behaves on a small test set shaped like real API data:</p>
<pre><code class="language-text">Study                          Ages    Site       age=28   age=28, Texas
Sleep and Metabolism in PCOS   18-40   Texas      kept     kept
Teen PCOS Study                12-17   Colorado   removed  removed
Online Support Program         18+     (none)     kept     removed
</code></pre>
<p>Notice the last row: some studies are fully remote and list no physical site, so a strict state filter hides them. A nice improvement is to keep those and label them "remote or site not listed".</p>
<h2>Step 4: Rank by meaning with embeddings</h2>
<p>Now the AI part. A person might write <em>"I have irregular cycles and I'm worried about my blood sugar"</em>. A study might describe itself as <em>"insulin resistance and metabolic outcomes in PCOS"</em>. They share almost no words, but they clearly belong together.</p>
<p>Just like the "Milky Way" example in my national parks post, <strong>embeddings</strong> capture meaning instead of exact words, so these two texts end up close together.</p>
<pre><code class="language-python">from sentence_transformers import SentenceTransformer
model = SentenceTransformer("all-MiniLM-L6-v2")


def rank_studies(candidates, user_text, top_n=5):
    if candidates.empty:
        return candidates

    docs = (
        candidates["title"] + ". " + candidates["summary"]
    ).tolist()

    doc_emb = model.encode(
        docs,
        normalize_embeddings=True
    )

    query_emb = model.encode(
        [user_text],
        normalize_embeddings=True
    )

    scores = (doc_emb @ query_emb.T).flatten()

    return (
        candidates
        .assign(score=scores)
        .sort_values("score", ascending=False)
        .head(top_n)
    )


user_text = "I have irregular cycles and I'm worried about my blood sugar"

top = rank_studies(
    candidates,
    user_text
)

print(
    top[
        ["nct_id", "title", "score"]
    ]
)
</code></pre>
<h2>Step 5: Explain each study in plain language</h2>
<p>Eligibility criteria are written for researchers. A language model can translate them into everyday words. But in healthcare, grounding is not optional. The model must only use the study's official text, must not guess whether the person qualifies, and should turn uncertainty into questions for their doctor.</p>
<pre><code class="language-python">import anthropic

client = anthropic.Anthropic()
# reads the ANTHROPIC_API_KEY environment variable


def explain_study(study, user_text):
    prompt = f"""
    You help people understand clinical studies in plain language.

    A person described their situation as:
    "{user_text}"

    Official study information:

    Title:
    {study.title}

    Summary:
    {study.summary}

    Eligibility criteria:
    {study.criteria}

    Write three short parts:

    1. What this study is about, in simple words.
    2. The main eligibility requirements, in simple words.
    3. Two or three questions this person could ask their doctor or the study team.

    Rules:
    use ONLY the official information above.
    Do not say whether the person qualifies.
    Do not give medical advice.
    If something is unclear, say so.
    """

    message = client.messages.create(
        model="claude-sonnet-5-5",
        max_tokens=600,
        messages=[
            {
                "role": "user",
                "content": prompt
            }
        ],
    )

    return message.content[0].text


for study in top.head(3).itertuples():
    print(study.title, study.url)
    print(explain_study(study, user_text))
</code></pre>
<blockquote>
<p><strong>Why this design matters:</strong> the code decides <em>which</em> studies to show, using strict rules and official data. The LLM only <em>explains</em> them. It never picks studies, never invents new ones, and never decides eligibility. That separation is what makes an AI tool trustworthy in a sensitive area like health.</p>
</blockquote>
<h2>Step 6: Build a simple web app</h2>
<p>Save the functions above in <code>trials.py</code>, then create <code>app.py</code>:</p>
<pre><code class="language-python">import streamlit as st
from trials import fetch_pcos_trials, to_table, basic_filter, rank_studies, explain_study

st.title("PCOS Research Study Finder")

st.caption(
    "For discovery only. This is not medical advice. "
    "Talk with your doctor before joining any study."
)


@st.cache_data(ttl=24 * 3600)
# download once a day, not on every click
def load_trials():
    return to_table(fetch_pcos_trials())


age = st.number_input(
    "Your age",
    min_value=12,
    max_value=80,
    value=28
)

state = st.text_input(
    "US state (optional, full name)",
    ""
)

user_text = st.text_area(
    "In your own words, what would you like help with?",
    "irregular cycles and questions about my metabolism"
)


if st.button("Find studies"):
    candidates = basic_filter(
        load_trials(),
        age,
        state or None
    )

    top = rank_studies(
        candidates,
        user_text
    )

    if top.empty:
        st.info(
            "No matching recruiting studies right now. "
            "Try removing the state filter."
        )

    for study in top.itertuples():
        st.subheader(study.title)

        st.markdown(
            f"View official study page"
        )

        with st.spinner("Explaining in plain language..."):
            st.markdown(
                explain_study(study, user_text)
            )
</code></pre>
<p>Run it with <code>streamlit run app.py</code>.</p>
<h2>Step 7: Designing responsibly</h2>
<p>Building for health is different from recommending movies. These are the principles I followed, and I'd suggest them for any similar project:</p>
<ul>
<li><p><strong>Discovery, not decisions.</strong> The tool points people to studies and questions. The doctor and the study team make every decision.</p>
</li>
<li><p><strong>Rules before AI.</strong> Hard requirements like age are applied by plain code, so the AI can't override them.</p>
</li>
<li><p><strong>Grounded explanations.</strong> The LLM only sees official study text and is told not to judge eligibility.</p>
</li>
<li><p><strong>Privacy by design.</strong> The app doesn't store what people type. If you deploy it, avoid logging user text, since health details are sensitive. Also be aware that user text is sent to the LLM provider, so say so clearly in your app.</p>
</li>
<li><p><strong>Always link the source.</strong> Every result links to the official ClinicalTrials.gov page, where the full details and contacts live.</p>
</li>
<li><p><strong>Respectful language.</strong> PCOS is a medical condition, not a personal failing. The app never mentions weight or blame, and it speaks to the person, not about them.</p>
</li>
</ul>
<h2>How to measure if it works</h2>
<p>For this kind of tool, the most important measure is <strong>safety</strong>: no result should ever break a hard rule like the age range, so write automated tests for the filters. Beyond that, check <strong>relevance</strong> (do people find the top studies worth reading?) and <strong>clarity</strong> (can a non-expert understand the explanation?). Asking a few people with PCOS to try the app and share honest feedback is more valuable than any metric.</p>
<h2>Ideas to take it further</h2>
<p>You could use the API's location filter to search by distance from the user's city, show which studies are remote-friendly, support other languages, or send a gentle alert when a new matching study starts recruiting. The same design also works for other conditions where research needs more participants, such as endometriosis.</p>
<h2>Conclusion</h2>
<p>We built a tool that turns hundreds of technical research listings into a short, understandable list of studies worth discussing with a doctor. It uses live public data, strict rules for safety, embeddings to understand what people mean, and an LLM that explains without inventing.</p>
<p>Most of all, it shows that recommendation systems can do more than suggest the next video to watch. Pointed at the right problem, they can connect people with research that may one day improve care for millions.</p>
<p><em>Data source: ClinicalTrials.gov, U.S. National Library of Medicine. This project is for educational purposes only and is not medical advice.</em></p>
]]></content:encoded></item><item><title><![CDATA[Building an AI recommender for national parks]]></title><description><![CDATA[How to match travelers with the right park using Python, text similarity, embeddings and a language model, with working code you can run today.
Tutorial for beginners and intermediate Python developer]]></description><link>https://rsp180002.hashnode.dev/building-an-ai-recommender-for-national-parks</link><guid isPermaLink="true">https://rsp180002.hashnode.dev/building-an-ai-recommender-for-national-parks</guid><dc:creator><![CDATA[Rushi]]></dc:creator><pubDate>Tue, 29 Sep 2026 03:32:19 GMT</pubDate><content:encoded><![CDATA[<p>How to match travelers with the right park using Python, text similarity, embeddings and a language model, with working code you can run today.</p>
<p>Tutorial for beginners and intermediate Python developers. About 12 minutes to read.</p>
<p>There are 63 national parks in the United States, and hundreds more around the world. Most travelers only hear about the same five or six. A recommendation system can change that by asking a simple question: <em>what do you actually enjoy outdoors?</em> Then it suggests parks you may never have considered, like a quiet desert under dark skies instead of a crowded valley in peak season.</p>
<p>In this article we will build one step by step. We start with a small dataset, add classic recommendation techniques, then use AI embeddings and a large language model to make the results smarter and easier to understand.</p>
<h2>What is a recommendation system?</h2>
<p>A recommendation system predicts which items a person is likely to enjoy. Netflix recommends shows, Spotify recommends songs, and our system will recommend parks. There are three main approaches, and good systems usually combine them.</p>
<p><strong>Content-based:</strong> Compares the features of parks, such as waterfalls, desert or wildlife, with what the user likes.</p>
<p><strong>Collaborative:</strong> Learns from other travelers: people who loved Zion also loved the Grand Canyon.</p>
<p><strong>AI and hybrid:</strong> Uses embeddings to understand meaning, and a language model to explain each pick in plain words.</p>
<h2>The plan</h2>
<ol>
<li><p>Build a dataset that describes each park.</p>
</li>
<li><p>Recommend similar parks with TF-IDF and cosine similarity.</p>
</li>
<li><p>Recommend parks from a user's own words, with filters for season and crowds.</p>
</li>
<li><p>Add collaborative filtering from traveler ratings.</p>
</li>
<li><p>Upgrade to semantic search with AI embeddings.</p>
</li>
<li><p>Use a language model to explain the recommendations.</p>
</li>
<li><p>Turn it into a simple web app.</p>
</li>
</ol>
<p>Install the libraries first:</p>
<pre><code class="language-plaintext">pip install pandas scikit-learn sentence-transformers anthropic streamlit
</code></pre>
<h2>Step 1: Create the park dataset</h2>
<p>Every recommender starts with data. For this tutorial we write a small dataset by hand. Each park has a short description, the best season to visit, and a crowd level from 1 (very quiet) to 5 (very busy).</p>
<pre><code class="language-plaintext">import pandas as pd

parks = pd.DataFrame([
    {"name": "Yellowstone", "state": "WY", "best_season": "summer", "crowd_level": 5,
     "description": "geysers hot springs wildlife bison wolves hiking camping volcanic plateau"},
    {"name": "Yosemite", "state": "CA", "best_season": "spring", "crowd_level": 5,
     "description": "granite cliffs waterfalls rock climbing giant sequoias hiking valley"},
    {"name": "Grand Canyon", "state": "AZ", "best_season": "spring", "crowd_level": 5,
     "description": "canyon views rim hiking rafting Colorado River desert sunsets"},
    {"name": "Zion", "state": "UT", "best_season": "spring", "crowd_level": 5,
     "description": "slot canyons river hiking red rock cliffs desert canyoneering"},
    {"name": "Acadia", "state": "ME", "best_season": "fall", "crowd_level": 4,
     "description": "rocky coastline ocean sunrise mountains biking carriage roads fall colors"},
    {"name": "Great Smoky Mountains", "state": "TN/NC", "best_season": "fall", "crowd_level": 5,
     "description": "misty forests wildflowers black bears waterfalls hiking fall colors"},
    {"name": "Glacier", "state": "MT", "best_season": "summer", "crowd_level": 4,
     "description": "alpine lakes glaciers mountains scenic drive mountain goats hiking"},
    {"name": "Everglades", "state": "FL", "best_season": "winter", "crowd_level": 2,
     "description": "wetlands alligators kayaking birdwatching mangroves wildlife"},
    {"name": "Olympic", "state": "WA", "best_season": "summer", "crowd_level": 3,
     "description": "temperate rainforest moss beaches ocean mountains hiking quiet forests"},
    {"name": "Great Basin", "state": "NV", "best_season": "summer", "crowd_level": 1,
     "description": "dark skies stargazing ancient bristlecone pines caves quiet mountains"},
    {"name": "Big Bend", "state": "TX", "best_season": "winter", "crowd_level": 1,
     "description": "desert river canyons hot springs stargazing dark skies remote hiking"},
    {"name": "Denali", "state": "AK", "best_season": "summer", "crowd_level": 2,
     "description": "tundra wildlife grizzly bears moose mountains remote wilderness"},
])
</code></pre>
<p>For a real project, use the free <a href="https://www.nps.gov/subjects/developer/index.htm">National Park Service API</a>. It returns official descriptions, activities, topics and alerts for every park, which gives your recommender much richer data.</p>
<h2>Step 2: Find similar parks with TF-IDF</h2>
<p>Computers cannot compare sentences directly, so we turn each description into numbers. <strong>TF-IDF</strong> gives each word a weight: words that are common in one park but rare across all parks get a high score. "Stargazing" says more about a park than "hiking", because almost every park has hiking.</p>
<p>Then <strong>cosine similarity</strong> measures how close two parks are, from 0 (nothing in common) to 1 (identical).</p>
<pre><code class="language-plaintext">from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.metrics.pairwise import cosine_similarity

vectorizer = TfidfVectorizer(stop_words="english")
park_matrix = vectorizer.fit_transform(parks["description"])
similarity = cosine_similarity(park_matrix)

def similar_parks(park_name, top_n=3):
    idx = parks.index[parks["name"] == park_name][0]
    result = parks.assign(score=similarity[idx]).drop(index=idx)
    return result.sort_values("score", ascending=False).head(top_n)[["name", "state", "score"]]

print(similar_parks("Zion"))
</code></pre>
<pre><code class="language-plaintext">        name state  score
    Big Bend    TX   0.32
    Yosemite    CA   0.24
Grand Canyon    AZ   0.20
</code></pre>
<p>If you loved Zion, the system suggests Big Bend: another desert park with canyons and a river, but far fewer visitors.</p>
<h2>Step 3: Recommend from the user's own words</h2>
<p>Travelers do not always have a favorite park. They have wishes. We can put their wishes through the same vectorizer and compare them to every park, then apply filters for season and crowd level.</p>
<pre><code class="language-plaintext">def recommend_for_user(preferences, season=None, max_crowd=5, top_n=3):
    query_vec = vectorizer.transform([preferences])
    scores = cosine_similarity(query_vec, park_matrix).flatten()
    result = parks.assign(score=scores)
    if season:
        result = result[result["best_season"] == season]
    result = result[(result["crowd_level"] &lt;= max_crowd) &amp; (result["score"] &gt; 0)]
    return result.sort_values("score", ascending=False).head(top_n)

picks = recommend_for_user("quiet hiking, stargazing and dark skies", max_crowd=2)
print(picks[["name", "state", "score"]])
</code></pre>
<pre><code class="language-plaintext">       name state  score
Great Basin    NV   0.61
   Big Bend    TX   0.54
</code></pre>
<h2>Step 4: Learn from other travelers</h2>
<p>Content-based filtering only knows what is written in the descriptions. <strong>Collaborative filtering</strong> learns hidden patterns from behavior. If many people rate Zion and the Grand Canyon highly, those parks are related, even if their descriptions look different.</p>
<pre><code class="language-plaintext">ratings = pd.DataFrame({
    "user":   ["ana", "ana", "ana", "ben", "ben", "cara", "cara", "cara", "dev", "dev"],
    "park":   ["Zion", "Big Bend", "Grand Canyon", "Zion", "Grand Canyon",
               "Acadia", "Olympic", "Glacier", "Olympic", "Big Bend"],
    "rating": [5, 4, 5, 4, 5, 5, 4, 5, 5, 3],
})

# Rows are users, columns are parks, missing ratings become 0
user_park = ratings.pivot_table(index="user", columns="park", values="rating").fillna(0)

# Compare parks by how the same users rated them
park_sim = pd.DataFrame(cosine_similarity(user_park.T),
                        index=user_park.columns, columns=user_park.columns)

def fans_also_liked(park, top_n=3):
    return park_sim[park].drop(park).sort_values(ascending=False).head(top_n)

print(fans_also_liked("Zion"))
</code></pre>
<pre><code class="language-plaintext">Grand Canyon    0.99
Big Bend        0.62
Acadia          0.00
</code></pre>
<p>With real data you would have thousands of users. At that scale, libraries such as <code>surprise</code> or <code>implicit</code> use matrix factorization, which handles large, sparse rating tables much better.</p>
<h2>Step 5: Add AI with semantic embeddings</h2>
<p>TF-IDF has a weakness: it only matches exact words. If a user types "I want to see the Milky Way", TF-IDF finds nothing, because no description contains "Milky Way". A person would instantly know this means stargazing.</p>
<p><strong>Embeddings</strong> fix this. A neural network turns each text into a list of numbers that captures its <em>meaning</em>. Texts with similar meanings end up close together, even when they share no words.</p>
<pre><code class="language-plaintext">from sentence_transformers import SentenceTransformer

model = SentenceTransformer("all-MiniLM-L6-v2")   # small, fast, free, runs on a laptop
park_embeddings = model.encode(parks["description"].tolist(), normalize_embeddings=True)

def semantic_search(query, top_n=3, max_crowd=5):
    query_emb = model.encode([query], normalize_embeddings=True)
    scores = (park_embeddings @ query_emb.T).flatten()   # cosine similarity
    result = parks.assign(score=scores)
    result = result[result["crowd_level"] &lt;= max_crowd]
    return result.sort_values("score", ascending=False).head(top_n)

picks = semantic_search("I want to see the Milky Way far away from city lights")
print(picks[["name", "state", "score"]])
</code></pre>
<p>Now the search understands meaning, and dark-sky parks like Great Basin and Big Bend rise to the top. For hundreds of thousands of items you can store embeddings in a vector database such as FAISS, Chroma or pgvector.</p>
<h3>Combining the signals (hybrid)</h3>
<p>The strongest systems blend several scores. A simple and effective formula is a weighted sum, for example <code>0.6 × semantic score + 0.3 × collaborative score + 0.1 × (1 − crowd level / 5)</code>. Tune the weights by testing what your users actually click and save.</p>
<h2>Step 6: Explain the picks with a language model</h2>
<p>People trust recommendations more when they understand <em>why</em>. A large language model can turn a list of scores into a friendly, personal explanation. Here is an example with the Anthropic Python SDK. Any modern LLM API works the same way.</p>
<pre><code class="language-plaintext">import anthropic

client = anthropic.Anthropic()   # reads the ANTHROPIC_API_KEY environment variable

def explain_picks(user_wishes, picks):
    park_list = "\n".join(
        f"- {row.name} ({row.state}), best in {row.best_season}: {row.description}"
        for row in picks.itertuples()
    )
    prompt = f"""A traveler said: "{user_wishes}"

Our recommender chose these national parks:
{park_list}

For each park, write two friendly sentences explaining why it fits this traveler.
Only use the facts given above. Do not invent details."""

    message = client.messages.create(
        model="claude-sonnet-5-5",
        max_tokens=500,
        messages=[{"role": "user", "content": prompt}],
    )
    return message.content[0].text

wishes = "I want to see the Milky Way far away from city lights"
print(explain_picks(wishes, semantic_search(wishes)))
</code></pre>
<p>Notice the design: the recommender chooses the parks, and the language model only explains them. This keeps results consistent and stops the model from inventing parks or facts. This pattern is often called retrieval-augmented generation (RAG).</p>
<h2>Step 7: Turn it into a web app</h2>
<p>With Streamlit you can give your recommender a simple interface in a few lines. Save this as <code>app.py</code> together with the code above, then run <code>streamlit run app.py</code>.</p>
<pre><code class="language-plaintext">import streamlit as st

st.title("Find your next national park")

wishes = st.text_area("What do you love outdoors?", "quiet forests, waterfalls and wildlife")
max_crowd = st.slider("Maximum crowd level", 1, 5, 5)

if st.button("Show recommendations"):
    picks = semantic_search(wishes, max_crowd=max_crowd)
    for row in picks.itertuples():
        st.subheader(f"{row.name}, {row.state}")
        st.caption(f"Best season: {row.best_season} | Crowd level: {row.crowd_level}/5")
        st.write(row.description)
    with st.spinner("Writing explanations..."):
        st.markdown(explain_picks(wishes, picks))
</code></pre>
<h2>How to measure if it works</h2>
<p>A recommender is only good if people like its suggestions. Common ways to measure that include <strong>precision at k</strong> (how many of the top k picks the user actually saved or visited), <strong>diversity</strong> (are the picks all the same kind of park?), and <strong>novelty</strong> (does it suggest parks beyond the famous ones?). Also plan for the <strong>cold start problem</strong>: new users have no ratings, so start with content-based and semantic search, and blend in collaborative filtering as ratings arrive.</p>
<h2>Ideas to take it further</h2>
<p>Once the basics work, you can add live weather and park alerts from the NPS API, distance from the user's home city, accessibility information, trip length, or photo-based search where users upload a picture of a landscape they love. Each new signal is one more feature your hybrid score can use.</p>
<h2>Conclusion</h2>
<p>We started with twelve short descriptions and ended with an AI-powered system that understands what travelers mean, learns from other visitors, and explains its choices in plain language. The same pattern works for hotels, hiking trails, books or restaurants. The best part of a park recommender is its purpose: helping people discover quiet, wonderful places they would never have found on their own.</p>
<p>Try the code, swap in real data from the NPS API, and share which park your recommender suggested for you.</p>
]]></content:encoded></item></channel></rss>