Developer productivity (internal) Managed Custom AI Agents AI Product Engineering Leanware · May 25, 2026

Engineering Metrics Dashboard: Tracking Estimate Accuracy

Leanware built an engineering metrics dashboard that pulls GitHub commit history into one view, then goes past standard output metrics to flag exactly where task estimates run long or short before the next planning cycle.

Geography
Internal
Stage
Internal Leanware build
Team
1 senior + 1 mid full-stack engineer, 1 product designer, 1 product owner

The situation

Most engineering metrics dashboards stop at output: commits, PRs merged, lines shipped. What they don't show is whether the team's estimates were any good, and that gap between estimated and actual time almost never turns into a metric anyone can act on. Leanware felt the same friction internally, building dashboards for output while the estimate vs actual data sat unreviewed in commit logs.

CodiQ started as an internal build to close that gap. The brief was an AI powered tool, owned by Leanware, that pulled commit history straight from a project's GitHub repository, compared it against the estimate committed to at scoping, and turned the difference into a clear pattern of over and under-estimation the team could feed into the next round of planning, an engineering metrics dashboard built in-house rather than bought off the shelf. 

What we built

The engagement ran the client framing in reverse: Leanware was both the team building the tool and the team using it. The build was staffed like any other AI Product Engineering project, with one senior and one mid full-stack engineer, a product designer, and a product owner. CodiQ connects to a project's GitHub repository, pulls the commit history, and pairs it with the task-level estimates that scoped the work in the first place. OpenAI runs the analysis: which tasks ran long, which finished early, and which patterns repeat across a single engineer or an entire team. That's the layer a standard output-metrics dashboard skips, and the layer CodiQ was built to surface. 

The user journey stays short by design:
- Connect the repository Select the scope window
- Run the analysis Review the AI verdict on estimation accuracy
- Under the hood, the stack is React on the frontend, Python and Django on the backend, Cloud - Run for compute
- Cloud Run Functions for the AI-call layer, PubSub for event handling, PostgreSQL for storage, and the OpenAI API for the analysis itself.
-CodiQ also doubled as a proof point during beta testing, a working demonstration that Leanware's AI fluency is operational rather than theoretical.

Outcome

  • AI analysis of GitHub commit history paired with task estimates running in beta

    Source ↗
  • Estimation-pattern insight delivered to Leanware engineers as the first cohort of users

    Source ↗

CodiQ landed inside the firm as a real working tool, not a prototype that stalled after the demo. The beta gave Leanware engineers data-driven feedback on their own estimation patterns instead of a gut feeling about who runs long. The planned expansions (advanced models, customizable dashboards, multi-VCS integration) are scoped against that working baseline rather than a green-field redesign, which keeps the roadmap grounded in what beta users actually asked for. 

Engagement FAQ

What is an engineering metrics dashboard, and how is CodiQ different from a standard one?

A typical engineering metrics dashboard tracks output commits, PRs, deploy counts. CodiQ pulls the same GitHub commit history but pairs it with the original task-level estimate, then uses OpenAI to flag where the two diverge. It's an output dashboard plus an estimate-accuracy layer most tools don't cover. 

What are developer productivity metrics based on GitHub commit history?

They come from comparing what a team estimated a task would take against what the commit history shows it actually took. CodiQ builds this by pulling a project's commit history and pairing it with the task-level estimate that scoped the work, then surfacing where the two diverge. 

How does an AI tool measure estimate accuracy from commit history?

CodiQ connects to a project's GitHub repository, pulls the commit history for a chosen scope window, and runs it through OpenAI alongside the original task estimates. The output flags which tasks ran long, which finished early, and which patterns repeat across an engineer or a team. 

Is CodiQ a DORA metrics dashboard?

No, DORA metrics track deployment frequency, lead time for changes, change failure rate, and time to restore service. CodiQ is scoped narrower: estimate accuracy from commit history versus the original task estimate. The two are complementary rather than the same thing, and multi-metric dashboards (DORA included) are part of CodiQ's planned expansion. 

What tech stack works well for building an internal AI analytics tool?

CodiQ runs on React for the frontend, Python and Django for the backend, Cloud Run for compute, Cloud Run Functions for the AI-call layer, PubSub for event handling, PostgreSQL for storage, and the OpenAI API for the analysis itself. 

Can an estimate-accuracy tool integrate with version control systems besides GitHub?

CodiQ currently connects to GitHub. Multi-VCS integration is on the roadmap as a planned expansion, alongside advanced models and customizable dashboards, scoped against the working beta rather than a rebuild.

What team does it take to build a tool like this internally?

 CodiQ shipped with one senior and one mid full-stack engineer, a product designer, and a product owner, run the same way Leanware staffs an external AI Product Engineering build.

react python django openai github developer tools internal

Related cases

MVP shipped within scope across iOS, Android, and web

Asana Rebel

Wellness app development for a global fitness brand: UnVape, a cross-platform quit-vaping app, shipped production-ready on iOS, Android, and web from one codebase in a few weeks.

Managed Custom AI Agents Read case study
Github Repository

Leanware

Generative AI for product managers, in practice: PRD Agent turns a short brief into a structured, reviewable PRD in minutes. Open source, so any team can run or extend it.

Managed Custom AI Agents Read case study
READY?

Stop managing operations. Let the system run them.

Show us the workflow that's eating your week. We will map it, show you what AI can automate, and tell you what we will run for you.

Tell us what you are trying to solve. We will map your workflows and show you exactly what AI can automate, and what we will run for you.