JobScout-AI
An agentic AI job search pipeline powered by Groq/Gemini, LangChain, and Chroma DB.
Automated Parsing
Extracts text from candidate CVs (.pdf, .docx) using PyPDF and processes it efficiently.
Learn moreReal-time Search
Executes intelligent, real-time job searches via the Tavily API based on LLM-generated queries.
Learn moreOverview
JobScout-AI is a comprehensive, agentic AI-driven system designed to completely automate and supercharge the job search process. Rather than relying on rigid keyword matching, JobScout-AI understands the semantic depth of a candidate's profile to discover highly relevant, real-time job postings.
The Core Problem: Traditional job hunting is tedious. Candidates spend hours scanning through job boards, filtering out irrelevant roles, and trying to match their unique skills against vague job descriptions. Standard keyword search often misses excellent opportunities hidden behind different terminology.
The Solution: By leveraging state-of-the-art Vector Language Models (VLMs) and embedding databases, JobScout-AI transforms a simple PDF resume into a dynamic, searchable vector. The AI agent acts as a personal recruiter: it reads your resume, understands your career trajectory, generates intelligent queries tailored to your profile, and scrapes the web (Indeed, LinkedIn) to find roles that perfectly align with your experience and aspirations.
Installation
Follow these steps to set up JobScout-AI locally on your machine.
Dependencies (requirements.txt)
Create a file named requirements.txt in your project root and copy-paste the following exact libraries that the project needs:
pymupdf
google-genai
chromadb
langchain-chroma
langchain-huggingface
langchain-text-splitters
python-jobspy
pandas
python-dotenv
sentence-transformers
Step-by-step Setup
# 1. Clone the repository
git clone https://github.com/yourusername/jobscout-ai.git
cd jobscout-ai
# 2. Create a virtual environment (highly recommended)
python -m venv venv
# On Windows use:
venv\Scripts\activate
# On Mac/Linux use:
source venv/bin/activate
# 3. Install the exact required packages
pip install -r requirements.txt
Configuration
To securely manage your API keys, the project uses a local .env file. Create a file named exactly .env in the root directory (jobscout-ai/) and add your keys:
GEMINI_API_KEY=your_gemini_api_key_here
Libraries & Downloads
-
Python Dotenv: Used to securely load environment variables from the `.env` file into your Python application.
bash
pip install python-dotenv
Usage
To run the full pipeline, place your target candidate CVs (PDFs or DOCX) into the
Data/inputs/ directory.
Then, simply execute the main orchestrator script:
python main.py
The application will process the CVs, generate queries, search for jobs, and save the
generated markdown reports into the Data/outputs/ directory.
1. Parser
The first critical step in the JobScout-AI pipeline is reliably extracting unstructured text from candidate CVs.
Overview
This module scans the Data/inputs/ directory for user resumes (supporting `.pdf` and `.docx` formats). It carefully reads the document layer by layer, extracting raw text while maintaining the original flow. Once the raw text is extracted, it intelligently chunks the document into smaller, semantically meaningful segments so that the VLM can process the data without exceeding context windows.
Libraries & Downloads
-
PyMuPDF: High-performance library for PDF rendering and text extraction.
bash
pip install PyMuPDF -
LangChain Text Splitters: Used specifically for the
RecursiveCharacterTextSplitterto chunk text efficiently.bashpip install langchain-text-splitters
2. VLM Analysis
Once the raw text is extracted, it is passed to a powerful Vector Language Model (VLM) for deep semantic structuring.
Overview
Raw text is messy. The VLM agent takes the chaotic chunks of resume text and uses a highly constrained prompt to restructure it into a strict JSON format. It categorizes the text into explicit fields like personal_info, skills, projects, and experience. This guarantees that the downstream embedding process has clean, normalized data to work with.
Libraries & Downloads
-
Google GenAI: The official Google SDK used to connect to the Gemini Flash-Lite model for ultra-fast, structured JSON generation.
bash
pip install google-genai
3. Chroma Store
The structured CV sections are converted into high-dimensional vectors and stored locally.
Overview
By embedding the resume data into a local ChromaDB instance, the system creates a "memory" of the candidate. This persistent vector store allows the agent to perform blazing-fast semantic similarity searches. When a job description is evaluated later, the system can instantly measure the mathematical distance between the candidate's embedded skills and the job's requirements.
Libraries & Downloads
-
ChromaDB: The open-source embedding database for building AI applications locally.
bash
pip install chromadb -
LangChain HuggingFace: Wrapper to easily use HuggingFace embedding models locally without external API calls.
bash
pip install langchain-huggingface -
Sentence Transformers: The underlying engine required to compute the dense vector embeddings locally.
bash
pip install sentence-transformers
4. Query Generation
The AI generates highly targeted job search queries based entirely on the parsed candidate profile.
Overview
Instead of relying on the user to guess the best search terms, the AI agent dynamically formulates complex, optimized search queries. It looks at the candidate's seniority, core technologies, and industry experience to generate strings like "Junior Python Backend Developer React". This ensures the scraping phase is highly effective.
Libraries & Downloads
-
Google GenAI: Used again to orchestrate the prompt logic that converts candidate JSON into optimized search terms.
bash
pip install google-genai
5. Job Search
The generated queries are executed against real-time job boards to fetch live opportunities.
Overview
Using the optimized queries, the system scrapes Indeed and LinkedIn simultaneously. It bypasses the need for expensive API keys by parsing the HTML of live job boards directly. It filters out stale jobs (e.g., older than 24 hours), handles pagination, and returns a clean, structured dataframe of fresh opportunities. These jobs are then scored against the candidate's ChromaDB vectors to find the perfect match.
Libraries & Downloads
-
Python JobSpy: A robust, anti-ban scraper designed specifically to extract job listings from LinkedIn, Indeed, Glassdoor, etc.
bash
pip install python-jobspy -
Pandas: The industry standard for data manipulation, used here to clean, filter, and structure the scraped job data.
bash
pip install pandas
Data Directory
The Data/ directory is used for local file storage and should not be committed
to version control.
Data/inputs/: Place candidate CVs here.Data/outputs/: Markdown reports are generated here.Data/chroma_db/: Persistent storage for the vector database.