Skip to content
Back to selected work

Case study / 2026

AI Resume Analyzer & Job-Description Matcher

An NLP application that scores how well a resume matches a job description and reports the exact skills to add.

Primary result

Instant match score + skill-gap report

  • Python
  • Scikit-learn
  • NLP
  • TF-IDF
  • Streamlit
AI Resume Analyzer & Job-Description Matcher2026 / shipped
RESUMEJOB DESCRIPTIONMATCH SCORE + SKILL GAP
TF-IDF
Vectorisation
Vectorisation
Cosine
Similarity
Similarity
PDF+TXT
Resume input
Resume input
NLP
Text pipeline
Text pipeline

Challenge

The problem to solve

Applicants rarely know why their resume does or doesn't fit a role. The goal was an instant, explainable match score plus a concrete skill-gap report.

Approach

Technical direction

Cleaned and tokenised resume and job-description text, vectorised both with TF-IDF (uni- and bi-grams), and computed cosine similarity — blended with skill-coverage against a curated tech vocabulary — for an intuitive match score. Built a Streamlit interface for upload-and-analyse.

Outcomes

Key outcomes

  1. 01Built an NLP matcher using TF-IDF and cosine similarity, blended with skill-coverage for an intuitive fit score.
  2. 02Engineered a text pipeline — cleaning, tokenisation and keyword extraction — over unstructured resume and job-description documents.
  3. 03Detects matched vs. missing skills against a curated technical vocabulary and surfaces the job description’s top keywords.
  4. 04Delivered a Streamlit app supporting PDF and text resume upload with an instant match score and skill-gap report.

Process

Process & architecture

01

Problem

Tailoring a resume to each role is guesswork. The project provides an objective match score and a specific list of skills to add before applying.

02

Inputs

A resume (uploaded as PDF or text) and a pasted job description. Text is extracted from PDFs and normalised for analysis.

03

Text preprocessing

Both documents are lowercased, stripped of noise and collapsed to clean tokens, preserving technical tokens like c++ and c#.

04

Matching approach

TF-IDF vectors (uni- and bi-grams, English stop-words removed) are compared with cosine similarity, then blended with skill-coverage so the score reflects how many of the role’s required skills the resume actually has.

05

Skill-gap analysis

A curated technical vocabulary is matched (word-boundary aware) against both texts to report skills the resume already covers versus those it is missing.

06

Interface

A Streamlit app returns an overall match score, matched vs. missing skills, and the top keywords from the job description — with a sample resume and JD preloaded to try instantly.

07

Future improvements

Sentence-transformer embeddings for semantic matching, section-aware parsing, and ranking one resume against many job descriptions.