CAT Used Price Advisor
A dealer-facing tool, powered by machine learning models, that prices more than 70,000 pieces of equipment worth over $12B.
A handful of projects I'm proud of, spanning equipment pricing, genomics, and plant breeding — the kind of work I keep coming back to when explaining what I actually do all day.
A dealer-facing tool, powered by machine learning models, that prices more than 70,000 pieces of equipment worth over $12B.
An index that separates real market price movement from quality differences, used by CAT Financial and senior leadership to sharpen pricing strategy across sale channels.
A forecasting model that predicts industry size five quarters out, using price-trend data from auction and retail markets along with consumer behavior.
AI-assisted engineering workflows spanning code generation, debugging, refactoring, and rapid prototyping, accelerating development and shortening delivery cycles.
An LLM-powered workflow that turns analytical outputs into one-page executive briefs, accelerating insight communication and decision-making.
KPI dashboards and decision support for SVPs, translating complex analyses into recommendations for product strategy and patent disputes.
An interpolation-based genomic coordinate conversion algorithm, served via an R Shiny app, that cuts conversion time by 90%.
Queryable BLAST databases built from public genomes, paired with a haplotype-based method for identifying gene-editing targets with greater statistical power — avoiding $750K in cost while widening research accessibility.
A parent-progeny-trio QC framework that delivers 100% data quality control with zero sample mix-ups.
Production pipelines processing 3TB of genomic data a day, cutting manual effort by 70% and making downstream analytics faster and more reliable.
A genomic feature repository in BigQuery that became the team's single source of truth for crop analytics, and the foundation for everything built on top of it since.
A proprietary imputation method that took usable data return from 30% to 92% per feature — a meaningful jump in how much of the data could actually be put to work.
Gaussian Mixture Models applied to improve true variant detection by 70% in genomic datasets, strengthening the data quality everything downstream depended on.
Insect-resistance genes mapped using mixed-model regression, then deployed into wheat varieties through hypothesis-driven experiments — resulting in products that use less pesticide and support more sustainable farming.
Sequencing-based bioinformatics methods to detect genome deletions and identify duplicate entries in public gene banks, cutting detection time by 60% and curation costs by 50%.
Principal component analysis and hierarchical clustering applied to identify a novel wheat population genetic group, contributing a new source of genetic variation to the research community.