Applied AI / LLM
Generative Unit Testing
A retrieval-augmented system that writes unit tests for a codebase. It indexes source files into a vector store, retrieves the most relevant context for a given target, and has an LLM draft tests grounded in how the surrounding code actually behaves — rather than guessing from a function signature alone. I led the prototype during a data-engineering internship in early 2024, when retrieval-augmented generation was still an emerging technique for grounding LLMs in private context.
Design
- Vector-store abstraction — an abstract storage interface with a concrete Pinecone implementation, so the retrieval backend can be swapped without touching the generation logic.
- Source parsing — a parser loads supported source files and splits them into chunks sized for retrieval.
- Pluggable embeddings — a dedicated embedder component wraps the embedding model behind a small interface.
- Grounded generation — retrieved context is passed to the LLM so generated tests reflect real usage, not just a signature.
Why RAG
Test quality depends on context. Feeding a model only the function under test produces shallow, often wrong assertions. Retrieving the surrounding code — callers, helpers, related types — gives the model the behaviour it needs to assert against, which is what makes the generated tests useful.
Where it fits now
The core problem is the same one today's coding assistants solve: give the model enough of the surrounding code to be correct, not just a signature. What has changed is the mechanism. In early 2024, a vector store was the practical way to pull in relevant context under tight token limits; modern tools often rely on much larger context windows and agents that navigate files directly. This project was an early, working take on that same goal — and the retrieval pattern it uses still holds up wherever a codebase is too large to hand a model all at once.