JobScout-AI Main Logo

JobScout-AI

An agentic AI job search pipeline powered by Groq/Gemini, LangChain, and Chroma DB.

Automated Parsing

Extracts text from candidate CVs (.pdf, .docx) using PyPDF and processes it efficiently.

Learn more

Real-time Search

Executes intelligent, real-time job searches via the Tavily API based on LLM-generated queries.

Learn more

Overview

JobScout-AI is a comprehensive, agentic AI-driven system designed to completely automate and supercharge the job search process. Rather than relying on rigid keyword matching, JobScout-AI understands the semantic depth of a candidate's profile to discover highly relevant, real-time job postings.

The Core Problem: Traditional job hunting is tedious. Candidates spend hours scanning through job boards, filtering out irrelevant roles, and trying to match their unique skills against vague job descriptions. Standard keyword search often misses excellent opportunities hidden behind different terminology.

The Solution: By leveraging state-of-the-art Vector Language Models (VLMs) and embedding databases, JobScout-AI transforms a simple PDF resume into a dynamic, searchable vector. The AI agent acts as a personal recruiter: it reads your resume, understands your career trajectory, generates intelligent queries tailored to your profile, and scrapes the web (Indeed, LinkedIn) to find roles that perfectly align with your experience and aspirations.

Installation

Follow these steps to set up JobScout-AI locally on your machine.

Dependencies (requirements.txt)

Create a file named requirements.txt in your project root and copy-paste the following exact libraries that the project needs:

pymupdf
google-genai
chromadb
langchain-chroma
langchain-huggingface
langchain-text-splitters
python-jobspy
pandas
python-dotenv
sentence-transformers

Step-by-step Setup

# 1. Clone the repository
git clone https://github.com/yourusername/jobscout-ai.git
cd jobscout-ai

# 2. Create a virtual environment (highly recommended)
python -m venv venv

# On Windows use: 
venv\Scripts\activate
# On Mac/Linux use: 
source venv/bin/activate 

# 3. Install the exact required packages
pip install -r requirements.txt

Configuration

To securely manage your API keys, the project uses a local .env file. Create a file named exactly .env in the root directory (jobscout-ai/) and add your keys:

GEMINI_API_KEY=your_gemini_api_key_here

Libraries & Downloads

  • Python Dotenv: Used to securely load environment variables from the `.env` file into your Python application.
    bash
    pip install python-dotenv

Usage

To run the full pipeline, place your target candidate CVs (PDFs or DOCX) into the Data/inputs/ directory.

Then, simply execute the main orchestrator script:

python main.py

The application will process the CVs, generate queries, search for jobs, and save the generated markdown reports into the Data/outputs/ directory.


1. Parser

The first critical step in the JobScout-AI pipeline is reliably extracting unstructured text from candidate CVs.

Overview

This module scans the Data/inputs/ directory for user resumes (supporting `.pdf` and `.docx` formats). It carefully reads the document layer by layer, extracting raw text while maintaining the original flow. Once the raw text is extracted, it intelligently chunks the document into smaller, semantically meaningful segments so that the VLM can process the data without exceeding context windows.

Libraries & Downloads

  • PyMuPDF: High-performance library for PDF rendering and text extraction.
    bash
    pip install PyMuPDF
  • LangChain Text Splitters: Used specifically for the RecursiveCharacterTextSplitter to chunk text efficiently.
    bash
    pip install langchain-text-splitters

2. VLM Analysis

Once the raw text is extracted, it is passed to a powerful Vector Language Model (VLM) for deep semantic structuring.

Overview

Raw text is messy. The VLM agent takes the chaotic chunks of resume text and uses a highly constrained prompt to restructure it into a strict JSON format. It categorizes the text into explicit fields like personal_info, skills, projects, and experience. This guarantees that the downstream embedding process has clean, normalized data to work with.

Libraries & Downloads

  • Google GenAI: The official Google SDK used to connect to the Gemini Flash-Lite model for ultra-fast, structured JSON generation.
    bash
    pip install google-genai

3. Chroma Store

The structured CV sections are converted into high-dimensional vectors and stored locally.

Overview

By embedding the resume data into a local ChromaDB instance, the system creates a "memory" of the candidate. This persistent vector store allows the agent to perform blazing-fast semantic similarity searches. When a job description is evaluated later, the system can instantly measure the mathematical distance between the candidate's embedded skills and the job's requirements.

Libraries & Downloads

  • ChromaDB: The open-source embedding database for building AI applications locally.
    bash
    pip install chromadb
  • LangChain HuggingFace: Wrapper to easily use HuggingFace embedding models locally without external API calls.
    bash
    pip install langchain-huggingface
  • Sentence Transformers: The underlying engine required to compute the dense vector embeddings locally.
    bash
    pip install sentence-transformers

4. Query Generation

The AI generates highly targeted job search queries based entirely on the parsed candidate profile.

Overview

Instead of relying on the user to guess the best search terms, the AI agent dynamically formulates complex, optimized search queries. It looks at the candidate's seniority, core technologies, and industry experience to generate strings like "Junior Python Backend Developer React". This ensures the scraping phase is highly effective.

Libraries & Downloads

  • Google GenAI: Used again to orchestrate the prompt logic that converts candidate JSON into optimized search terms.
    bash
    pip install google-genai

Data Directory

The Data/ directory is used for local file storage and should not be committed to version control.

  • Data/inputs/: Place candidate CVs here.
  • Data/outputs/: Markdown reports are generated here.
  • Data/chroma_db/: Persistent storage for the vector database.