Skip to content
All projects

GitHub · Case study

AI Engineering Assistant

Copilot-related engineering workflows grounded in repository context, with tool calling, streaming, and automated evaluation.

Illustration of a glowing crystalline sphere drawing in panels of code and streaming a response outward.
Illustration1 / 5
Illustration. The other images are screenshots of real public pages.
Role
Senior Software Engineer: AI-powered engineering workflows in Python.
Company
GitHub
Faster task completion reported by GitHub research
55%

Overview

GitHub Copilot-related engineering workflows built with Python and LLM APIs. Repository-aware retrieval supplies developer context, tool calling lets the model act on it, streaming keeps responses responsive, and automated evaluation tracks quality.

Problem

LLMs are only useful to engineers when they have the right repository context, respond quickly, and produce output whose quality can be measured rather than guessed.

Architecture

AI Engineering Assistant system designAssistant serviceModelDeveloperEditor · github.comAssistant APIAuth · streamingOrchestratorPython · promptsLLM APITool callingEvaluationAutomated scoringContext retrievalRepository-awareToolsFunctions it can callRepositorySource of contextstreampromptcallcontextsearchtest set
RequestAsync eventData read / writeMoving dots show which way data flows.
  1. Developer ↔ Assistant API: request, stream
  2. Assistant API ↔ Orchestrator: request
  3. Orchestrator ↔ LLM API: request, prompt
  4. LLM API ↔ Tools: request, call
  5. Orchestrator ↔ Context retrieval: data read / write, context
  6. Context retrieval ↔ Repository: data read / write, search
  7. Evaluation → Assistant API: async event, test set

What made it hard

  • Selecting relevant code and context from large repositories.
  • Fitting that context into model token limits.
  • Keeping perceived latency low with streaming responses.
  • Measuring quality with repeatable automated evaluation.

Solution

A Python workflow retrieves repository-aware context, assembles prompts, and calls LLM APIs with tool calling. Responses stream back to the developer, and automated evaluation scores outputs so changes can be compared objectively.

Skills used

  • Python
  • LLM APIs
  • GitHub Copilot
  • Repository-aware retrieval
  • Tool calling
  • Streaming
  • Automated evaluation

Impact

  • Contributed to Copilot-related workflows; GitHub research reported 55% faster task completion with Copilot.
  • Grounded model responses in real repository context.
  • Automated evaluation made quality changes measurable.

Decisions

  1. 01Retrieval quality matters more than prompt length.
  2. 02Stream early so developers see progress immediately.
  3. 03Run evaluation on every change instead of checking once by hand.