COMPUTER VISION · INSURANCE CLAIMS TRIAGE

Damage, or whole.
Flagged for human review in milliseconds.

ClaimSight is a binary vehicle-damage classifier for auto-insurance claim triage — from a deterministic ingestion pipeline through Grad-CAM explainability to a served FastAPI endpoint. It's a decision-support tool, not an autonomous adjuster: every flagged claim still routes to a human reviewer.

BUILD STATUS: training complete · ResNet50 selected (94.3% val. accuracy · ROC-AUC 0.986) · live demo up · API tested (34, 93% coverage), run locally
class: 00-damage · confidence: 0.89
class: 01-whole · confidence: 1.00
GRAD-CAM · REAL VALIDATION EXAMPLE
Grad-CAM heatmap overlay on a damaged vehicle door, correctly classified as 00-damage
PREDICTION00-damage
CONFIDENCE1.00
OUTCOMETrue Positive

Real overlay, unedited — coarse activation map, not a pixel-level segmentation mask (see §06).

01 · THE PROBLEM

Manual photo review doesn't scale with claim volume

Every submitted claim arrives with photos that need a first visual pass before an adjuster estimates repair cost. That pass is slow, inconsistent across reviewers, and adds latency before the customer hears anything back.

Triage, not adjudication

A lightweight binary pass routes obviously-intact vehicles away from the review queue and flags damaged ones with a visual cue — it never approves or denies a claim.

Asymmetric error cost

A damaged car missed by the classifier (false negative) is far more costly than a false alarm. Recall on the damage class is prioritized over raw accuracy.

Threshold as a business lever

The decision threshold is tuned against that cost asymmetry, not chosen purely to maximize a statistical score — and the trade-off is measured and documented in §05, not hidden.

02 · DATA ENGINEERING

One preprocess() function. Zero train/serve skew.

The same deterministic transform runs in training, evaluation, and the live API — so the tensor a model was trained on and the tensor it sees in production are built by identical code, not two implementations that drift apart.

01

Ingestion

Kaggle dataset pull (anujms/car-damage-detection) into a fixed folder layout — idempotent, re-runnable without duplication.

02

Deterministic preprocessing

BGR→RGB, aspect-preserving letterbox resize to 224×224, bilateral filter, CLAHE on the L-channel (LAB).

03

Augmentation

Albumentations transforms, applied strictly after the shared step, and only on the training split.

04

Normalization

ImageNet mean/std, fixed seeds, pinned requirements.txt for reproducibility.

05

Serving

The identical function is imported into the FastAPI handler — same bytes in, same tensor out.

ZERO-SKEW

src/preprocessing.py is imported by both src/dataset.py (training/evaluation) and app/main.py (serving). There is exactly one code path that turns an image into a model input — nothing is reimplemented for serving. Verified by an automated determinism test in tests/test_preprocessing.py.

Before/after comparison of three real claim photos through the preprocessing pipeline

Real output of src/preprocessing.preprocess() on three validation images — from notebooks/01_eda.ipynb.

03 · DATASET

2,300 labeled images, binary ground truth only

Source: Kaggle — anujms/car-damage-detection. Images are organized into training/ and validation/, each split into 00-damage and 01-whole.

2,300
Total images
3-channel RGB, variable native resolution
1,840 / 460
Train / Validation split
80% / 20% of the dataset
2
Classes
00-damage · 01-whole — no severity or part labels

Image count by split

Training — 1,840 (80%) Validation — 460 (20%)

Real counts from src/ingestion.dataset_summary() — not model output.

Per-class balance within each split

Split00-damage01-whole
Training920920
Validation230230

Perfectly balanced within each split — no class weighting needed in the loss function. Verified programmatically during ingestion (notebooks/01_eda.ipynb).

CONSTRAINT

No masks or bounding boxes are provided. This is a hard constraint that shapes the explainability approach (§06 — Grad-CAM as weak localization) and the future segmentation roadmap (§10).

04 · MODELING

Two backbones, two training phases, one clear winner

Transfer learning from ImageNet weights, compared head-to-head across two architectures with the identical two-phase training protocol, same seed, same data.

ResNet50 — IMAGENET1K_V2

Deeper residual backbone. Selected as the deployed model — see comparison below.

winnertorchvision.modelsfrozen → fine-tuned
val_accuracy 0.9435 · ROC-AUC 0.9858

EfficientNet-B0

Lighter compound-scaled backbone. Underperformed ResNet50 by ~5.6 points on this dataset size — plausibly because 1,840 training images favor ResNet50's more mature ImageNet1K_V2 transfer weights over EfficientNet's compound-scaling advantage, which tends to show up at larger data scale.

torchvision.modelsfrozen → fine-tuned
val_accuracy 0.8870 · ROC-AUC 0.9620
PHASE 1

Head training, backbone frozen

Only the classification head is trained (8 epochs, lr 1e-3); convolutional weights stay frozen to adapt quickly without disturbing pretrained features.

PHASE 2

Fine-tuning the last block

The final convolutional block is unfrozen and trained at a reduced learning rate (15 epochs, lr 1e-5), with ReduceLROnPlateau, early stopping (patience 5), and checkpoint selection by lowest validation loss.

Validation loss and accuracy curves for ResNet50 vs EfficientNet-B0 across both training phases

Real training curves from outputs/history/. Dotted line marks the Phase 2 (fine-tuning) boundary — both curves visibly improve once the last block unfreezes.

05 · RESULTS

94.3% validation accuracy, 0.986 ROC-AUC

Every number below is computed by src/evaluate.py on the 460-image held-out validation split for the deployed model (ResNet50) — raw JSON in outputs/metrics/metrics_resnet50.json.

All numbers on this page are real, computed on the model's actual held-out validation predictions — not estimates or placeholders. Reproduce with python -m src.evaluate --arch resnet50.

Confusion matrix (threshold = 0.50)

214
True Negative
16
False Positive
10
False Negative ⚑ costliest
220
True Positive

The FN cell is flagged because a missed damage claim is the most expensive error in this workflow (§01). 10 of 230 damaged vehicles were missed at the default threshold.

ROC curve for the damage class, AUC 0.9858

Metrics — damage class

0.957
Recall (priority)
220 / 230 damaged vehicles caught
0.932
Precision
16 false alarms out of 236 flagged
0.944
F1-score
harmonic mean, damage class
0.986
ROC-AUC
threshold-independent separability

Threshold tuning, business trade-off: lowering the decision threshold from 0.50 to 0.25 raises damage recall from 95.7% to 98.7% (missed claims drop from 10 to 3 of 230) at the cost of precision falling from 93.2% to 90.8% (false alarms rise from 16 to 23). In this workflow that trade is favorable: a false positive costs one extra human review; a false negative can wrongly close a legitimate claim.

Precision and recall on the damage class as a function of decision threshold
06 · EXPLAINABILITY

Grad-CAM shows where the model looked — not a mask

Grad-CAM (via pytorch-grad-cam) is computed on the last convolutional layer and rendered as overlays for both correct and incorrect predictions, across both classes — partly to check the model isn't relying on shortcuts like watermarks or backgrounds.

Grad-CAM (used here)True segmentation (roadmap)
OutputCoarse activation heatmapPixel-level mask
SourceGradients of a classifierDedicated annotation
RequiresNothing beyond class labelsMasks / polygons per image
MetricQualitative inspectionIoU / Dice

The dataset has no masks or bounding boxes, so Grad-CAM is used honestly as a weak, post-hoc localization signal for triage review — never described in this project as segmentation. True pixel-level segmentation is scoped as a separate future phase (§10), with its own dataset and metrics.

True positive Grad-CAM example, heatmap on vehicle door damage
TP · damage 0.89
False negative Grad-CAM example, heatmap on wheel/headlight instead of damage
FN · missed, 0.71
False positive Grad-CAM example, heatmap on reflective showroom trunk
FP · false alarm, 0.86
True negative Grad-CAM example, heatmap on intact vehicle body
TN · correct, 1.00

Honest shortcut-learning check: in several false negatives the heatmap concentrated on the wheel or headlight rather than the damaged panel, and in one false positive the model flagged a pristine showroom photo — likely reacting to reflections and a checkered dealership floor absent from most training images. Both are documented as real limitations (§10), not smoothed over.

07 · SERVING

A minimal, explicit API contract

Served with FastAPI + uvicorn, containerized via the included Dockerfile for any host. The interactive demo (streamlit_app.py) runs the same code live on Streamlit Community Cloud. The trained model itself is published to the Hugging Face Hub with a full model card documenting intended use and limitations.

GET/health
{
  "status": "ok",
  "model": "resnet50",
  "device": "cuda"
}
POST/predict
{
  "prediction": "00-damage",
  "confidence": 0.89,
  "probabilities": {
    "00-damage": 0.89,
    "01-whole": 0.11
  },
  "threshold_used": 0.5
}
08 · BUSINESS MODEL

How a triage API like this earns its keep

ClaimSight is a portfolio prototype, not a live commercial product — no payments are processed anywhere on this site. The pricing below illustrates how the underlying API contract (§07) maps directly onto a usage-based SaaS model that insurers and claims platforms already buy into today.

Free

$0
  • 50 predictions / month
  • Interactive demo access
  • Community support
  • Default 0.50 threshold
Try the demo ↓
Most common

Starter

$49 / month
  • 2,000 predictions / month
  • Configurable decision threshold
  • Email support, 48h SLA
  • Usage dashboard + API keys
Start trial →

Enterprise

Custom
  • Volume-based pricing
  • On-prem / VPC deployment
  • Fine-tuning on your claim archive
  • Dedicated support + uptime SLA
Contact sales →

Target customer: mid-size auto insurers, claims-management SaaS platforms, and body-shop networks that want an automated first-pass filter before a human adjuster ever opens a claim — cutting reviewer time on the ~49% of submissions (per this validation split) that show no visible damage at all.

09 · SKILL → EVIDENCE

Every claimed skill maps to a repository artifact

Every path below exists in the repository and is exercised by an automated test or a script that was actually run to produce the numbers on this page.

SkillEvidence
PyTorch transfer learning (2 backbones, 2 phases)src/train.py, src/model.py
Deterministic pipeline / zero-skewsrc/preprocessing.py → imported by src/dataset.py & app/main.py, verified in tests/test_preprocessing.py
Data augmentation (Albumentations)src/dataset.py
Business-oriented evaluation & threshold tuningsrc/evaluate.py, outputs/metrics/
Explainability (Grad-CAM)src/explain.py, outputs/gradcam/
Model serving (FastAPI, tested)app/main.py, tests/test_api.py (34 tests, 93% coverage)
CI rigor: unit/integration split, coverage gate, lintingtests/test_integration.py, pyproject.toml, .github/workflows/tests.yml
Reproducibility & pipeline engineeringrequirements.txt (pinned), fixed seeds, per-run metadata JSON
Model publication (Hugging Face Hub)scripts/publish_hf.py + generated model card
Containerized deploymentDockerfile
Free-tier live deploymentGitHub Pages, Hugging Face Hub, Streamlit Community Cloud
Interactive demo UI — livestreamlit_app.py
10 · LIMITATIONS & ROADMAP

What this system does not do — yet

Known limitations

  • Binary output only — no severity or damaged-part classification.
  • No masks or bounding boxes → Grad-CAM is weak localization only.
  • Modest dataset size (2,300 images, one source) → real domain-shift risk against a real insurer's photo distribution.
  • Observed false positives on reflective, studio-lit dealership photos — underrepresented in training data.
  • A small number of false negatives concentrated attention on wheels/headlights rather than the damaged panel.
  • Not validated as a repair-cost estimator — triage signal only, human review required.

Roadmap

  • Annotate a subset with CVAT / Label Studio, or adopt the CarDD dataset, for true pixel-level segmentation.
  • Train a U-Net via segmentation-models-pytorch and report IoU / Dice once real masks exist.
  • Extend to multi-class damage severity and part localization.
  • Expand training data with dealership/showroom photos to close the observed false-positive gap.
  • Add basic drift monitoring to the serving layer.