GitHub · Case study
AI Engineering Assistant
Copilot-related engineering workflows grounded in repository context, with tool calling, streaming, and automated evaluation.

- Role
- Senior Software Engineer: AI-powered engineering workflows in Python.
- Company
- GitHub
- Faster task completion reported by GitHub research
- 55%
Overview
GitHub Copilot-related engineering workflows built with Python and LLM APIs. Repository-aware retrieval supplies developer context, tool calling lets the model act on it, streaming keeps responses responsive, and automated evaluation tracks quality.
Problem
LLMs are only useful to engineers when they have the right repository context, respond quickly, and produce output whose quality can be measured rather than guessed.
Architecture
- Developer ↔ Assistant API: request, stream
- Assistant API ↔ Orchestrator: request
- Orchestrator ↔ LLM API: request, prompt
- LLM API ↔ Tools: request, call
- Orchestrator ↔ Context retrieval: data read / write, context
- Context retrieval ↔ Repository: data read / write, search
- Evaluation → Assistant API: async event, test set
What made it hard
- Selecting relevant code and context from large repositories.
- Fitting that context into model token limits.
- Keeping perceived latency low with streaming responses.
- Measuring quality with repeatable automated evaluation.
Solution
A Python workflow retrieves repository-aware context, assembles prompts, and calls LLM APIs with tool calling. Responses stream back to the developer, and automated evaluation scores outputs so changes can be compared objectively.
Skills used
- Python
- LLM APIs
- GitHub Copilot
- Repository-aware retrieval
- Tool calling
- Streaming
- Automated evaluation
Impact
- Contributed to Copilot-related workflows; GitHub research reported 55% faster task completion with Copilot.
- Grounded model responses in real repository context.
- Automated evaluation made quality changes measurable.
Decisions
- 01Retrieval quality matters more than prompt length.
- 02Stream early so developers see progress immediately.
- 03Run evaluation on every change instead of checking once by hand.