Files
motovaultpro/.claude/skills/mvp-research-frontier/SKILL.md
T
Eric GullicksonandClaude Fable 5 c239bb9347
Deploy to Staging / Build Images (push) Successful in 5m32s
Deploy to Staging / Deploy to Staging (push) Successful in 44s
Deploy to Staging / Verify Staging (push) Successful in 5s
Deploy to Staging / Notify Staging Ready (push) Successful in 4s
Deploy to Staging / Notify Staging Failure (push) Has been skipped
chore: replace AI skill library with 16 verified mvp-* skills
Removes the old planner/decision-critic/role-agents/domain-agents system
(step-injector scripts, stale scopes, sprint-era workflow) and replaces it
with 16 ground-truth-verified skills under .claude/skills/: change control,
debugging playbook, failure archaeology, architecture contract, domain
reference, OCR/Gemini pipeline, config and secrets, build and env, run and
operate, diagnostics (with tested scripts), validation and QA, docs and
writing, launch readiness, deploy-safety campaign, proof and analysis
toolkit, and research frontier. RULE 0/1/2, the temporal-contamination rule,
and the decision stress-test protocol are carried forward into the new
skills; the retired content remains in git history.

Co-Authored-By: Claude Fable 5 <[email protected]>
2026-07-09 20:21:42 -05:00

24 KiB

name, description
name description
mvp-research-frontier Open problems where MotoVaultPro can genuinely advance, plus the methodology for turning a hunch into an accepted result. Load when the issue tracker is empty and you need to seed the next work; when someone asks "what should we work on next", "what are the biggest open problems", "is this idea worth pursuing", or "how do I propose/validate a new improvement"; or when starting an experiment, evaluation harness, accuracy measurement, performance study, or multi-tenant/SaaS feasibility study. Trigger keywords - no open issues, roadmap, next steps, research, experiment, hypothesis, golden set, eval harness, accuracy of OCR/VIN extraction, multi-user readiness.

MotoVaultPro Research Frontier

Everything in this file is labeled candidate/open. Nothing here is a commitment, a roadmap, or owner-approved scope. This skill exists because the issue tracker was fully drained (zero open issues as of 2026-07-07) and the next work must be seeded deliberately, not by whichever bug happens next.

When to use / when NOT to use

Use this skill when:

  • You need to propose new work and there is no open issue telling you what to do.
  • You are about to start an experiment, measurement, or feasibility study and need the discipline that makes its result acceptable (Part B).
  • Someone asks whether one of the five frontier problems below is real or already solved.

Do NOT use this skill for:

  • Executing the issue/branch/PR workflow itself -- mvp-change-control.
  • Diagnosing a live failure -- mvp-debugging-playbook, then mvp-failure-archaeology to check whether the battle is already settled.
  • The CI/deploy-safety hole specifically -- mvp-deploy-safety-campaign is the executable campaign; Problem 3 below is only its extended end-state.
  • Root-cause evidence standards and benchmark discipline in depth -- mvp-proof-and-analysis-toolkit (this skill cross-references it, not replaces it).
  • What must be true before launch or public claims -- mvp-launch-readiness.
  • OCR/Gemini subsystem mechanics -- mvp-ocr-gemini-pipeline.
  • Running tests / what counts as validation evidence -- mvp-validation-and-qa.

Project reality this skill assumes (owner-confirmed 2026-07-07): development is done by AI sessions in this repo, the human owner reviews PRs, and end-to-end verification happens on staging via the PR deploy pipeline because there is no fully working local dev loop. Two owner non-negotiables apply to every experiment below: (1) no destructive database operations without a fresh backup (this includes the backend integration tests, which DROP TABLE ... CASCADE on the shared dev database -- see backend/src/features/vehicles/tests/integration/vehicles.integration.test.ts:38); (2) never hand-edit files on staging/prod servers -- rsync --delete on every deploy reverts them (.gitea/workflows/staging.yaml:117-118, production.yaml:108-109).


PART A -- Frontier problems (all candidate, verified against the repo 2026-07-07)

# Problem One-line asset Falsifiable milestone
1 LLM-extraction evaluation harness Real extraction pipeline + real receipts A model-swap PR shows a scorecard diff, not vibes
2 Type-safe DB boundary Recurring-bug corpus (#239/#241/#244; #243 reported but not confirmed) + shipped detector script Detector returns zero suspects AND a regression gate fails a planted bug
3 CI as a product asset Working per-PR staging deploy pipeline to extend A red PR cannot merge; staging never clobbered by unmerged PRs
4 Async OCR service Known sync-in-async defect with a measurable symptom /health p99 < 1s during saturated VIN decode, zero healthcheck restarts
5 Multi-user/SaaS readiness study Post-#206 UUID identity scoping already done Written go/no-go with each blocker proven by a test

Problem 1: LLM-extraction evaluation harness (candidate)

Why the current state fails. Receipt and VIN extraction accuracy has never been measured. All OCR tests mock Gemini and generate synthetic images inline with PIL (e.g. ocr/tests/test_receipt_extraction.py:269-291 builds a blank 100x100 image); there is no fixtures directory of real labeled images anywhere under ocr/tests/. Consequences already observed: the deployed model was changed to a preview model (GEMINI_MODEL: gemini-3-flash-preview, docker-compose.yml:210, overriding the gemini-2.5-flash code default in ocr/app/config.py) with no before/after accuracy evidence. Every future model or prompt change is currently judged by eyeball.

This project's asset. A complete, deployed extraction pipeline (ocr/app/extractors/, ocr/app/routers/{extract,decode}.py) with known output schemas to score against: fuel receipts produce totalAmount, fuelQuantity, pricePerUnit, fuelGrade (ocr/app/extractors/fuel_receipt.py:111-147); maintenance receipts produce serviceName, serviceDate, totalCost, shopName, laborCost, partsCost, odometerReading (ocr/app/extractors/maintenance_receipt_extractor.py:40-46); VIN decode produces year/make/model/trimLevel/engine/transmission (ocr/app/routers/decode.py, response model in ocr/app/models/schemas.py). Real usage exists: the owner's own uploaded documents persist in the data/documents/ volume (treat as PII -- get owner sign-off and strip identifying content before committing any image to git).

First three steps in this repo:

  1. Create ocr/tests/fixtures/golden/ with receipts/, maintenance/, vins/ subdirectories and one labels.json per subdirectory mapping filename to expected field values (start with N=10 per category; owner-supplied, sanitized images).
  2. Write ocr/tests/test_golden_extraction.py: for each labeled image, call the extractor and score field-level precision/recall (a field counts as correct only on exact match after normalization; report per-field and aggregate). Gate the whole module behind an env var (e.g. GOLDEN_EVAL=1 + pytest.mark.skipif) so the default cd ocr && python -m pytest run stays hermetic -- live Gemini credentials (the WIF chain, see mvp-ocr-gemini-pipeline) exist only in containers.
  3. Run the baseline on staging (the only environment with working credentials): shell into the OCR container and run the gated suite; commit the resulting scorecard JSON next to the fixtures with the model name and date. This number is the baseline every future prompt/model PR must diff against.

You have a result when a PR that changes GEMINI_MODEL, a prompt, or an extractor includes a before/after scorecard produced by the harness on the same golden set -- and a reviewer can reject it because a number went down. Falsified if, after building it, scores are so noisy across identical runs (LLM nondeterminism) that diffs are meaningless; that outcome itself is a publishable finding (then measure variance first and report mean-of-3).

Problem 2: Type-safe DB boundary -- eliminate the pg-numeric bug class (candidate)

Why the current state fails. node-postgres returns NUMERIC/DECIMAL (PostgreSQL type OID 1700) as strings. backend/src/core/config/database.ts:12 overrides only the DATE parser (OID 1082 -> plain string), so every repository mapper must hand-coerce numerics. This produced at least three shipped bugs (#239, #241, #244 -- #243 was reported for stations but the stations mappers already coerced via parseFloat, see mvp-failure-archaeology incident 12) with inconsistent coercion styles across repositories, and #244 showed a second failure shape: "enhanced" methods returning raw rows that bypass the mapper entirely. The convention is enforced by review only; nothing stops the next repository from reintroducing it.

This project's asset. The bug class is fully characterized (see mvp-failure-archaeology), and a working detector already ships: .claude/skills/mvp-diagnostics-and-logging/scripts/check-numeric-coercion.sh (harvests numeric columns from migration SQL, flags uncoerced row.<col> reads and raw-row returns in *.repository.ts; exit 0 = clean). There is also an unused codegen seed: cd backend && npm run schema:generate (backend/src/_system/schema/generate.ts).

First three steps in this repo:

  1. Run the detector and record the current suspect list as the baseline: bash .claude/skills/mvp-diagnostics-and-logging/scripts/check-numeric-coercion.sh.
  2. Decide the mechanism with an explicit trade-off writeup (stress-test per mvp-proof-and-analysis-toolkit). Three options, in ascending ambition: (a) shared mapNumeric() helper in backend/src/core/ that every mapper calls; (b) global OID 1700 parser in backend/src/core/config/database.ts -- the OID 1082 override at line 12 is the successful precedent, but weigh honestly: a global parser changes every consumer at once, and parseFloat on money columns trades string safety for float representation (audit currency columns before choosing); (c) schema-derived row types via schema:generate so uncoerced reads become compile errors. Record the decision and the rejected options in the issue.
  3. Fix all current suspects in one PR, then add the regression gate: wire the detector into backend/package.json lint (or an eslint no-restricted-syntax rule in backend/eslint.config.js banning return result.rows in repository files -- candidate, feasibility unproven) so a new violation fails locally and, once Problem 3 lands, in CI.

You have a result when (a) the detector script exits 0 on main, AND (b) a deliberately planted violation -- an uncoerced row.price read added on a scratch branch -- is caught by the automated gate without human review. Falsified if the global parser (option b) breaks an existing consumer that depended on string numerics: that is a finding, not a failure; document it in mvp-failure-archaeology and fall back to (a).

Problem 3: CI as a product asset (candidate -- end-state extension of the campaign)

Prerequisite: read and follow mvp-deploy-safety-campaign. This entry only records the frontier beyond that campaign's scope. Do not start here first.

Why the current state fails. CI gates nothing beyond build + boot (canonical statement: mvp-validation-and-qa section 1). Verified sharp edges: staging.yaml triggers on pull_request (line 12) and every PR build tags and pushes :latest (lines 58-98), so an unmerged PR overwrites both the shared staging environment and the very tag that production.yaml deploys by default (image_tag default 'latest', production.yaml:17). The mobile+desktop hard requirement is validated nowhere: the single Cypress spec (frontend/cypress/e2e/stations.cy.ts) has no config and cannot run.

This project's asset. A working self-hosted pipeline that already builds, deploys, and health-verifies a full stack per PR -- most projects have to build that from scratch; here only the gates and isolation are missing.

First three steps in this repo (after the campaign's own steps):

  1. Split image tagging in .gitea/workflows/staging.yaml: PR builds push only :{shortsha} (and optionally :pr-{n}); :latest moves only on push to main. This alone ends the unmerged-code-to-prod path.
  2. Add a test job ahead of build: backend unit tests (cd backend && npm test -- --forceExit, excluding tests/integration unless the job provisions an ephemeral postgres+redis -- never point CI tests at a shared database; they DROP TABLE CASCADE), backend/frontend npm run lint and npm run type-check, and cd ocr && python -m pytest. Red job blocks merge.
  3. Add a post-deploy viewport smoke job: Playwright (dep already in root package.json, currently dead -- no config exists) hitting https://staging.motovaultpro.com at 375x812 and 1280x800, asserting login page renders and no console errors. This is the first automated teeth behind the mobile+desktop rule.

You have a result when (a) a PR containing a deliberately failing unit test cannot be merged (workflow red, branch protection enforcing it), AND (b) after a PR build completes, docker manifest inspect (or registry UI) shows :latest unchanged -- staging mainline is never clobbered by an unmerged PR. Per-PR ephemeral environments or a namespaced staging are the stretch goal, candidate only; the milestone above does not require them.

Problem 4: Async OCR service -- Gemini off the event loop (candidate)

Why the current state fails. ocr/app/routers/decode.py calls the synchronous, unbounded _gemini_engine.decode_vin(vin) directly inside an async def handler (same pattern for sync receipt paths in routers/extract.py); only the manual-PDF path uses run_in_executor (extract.py:472, jobs.py:131). A Search-grounded VIN decode has been observed taking 34s+ (issue #230 history), during which the entire uvicorn event loop -- including GET /health (ocr/app/main.py:44) -- is blocked. The Docker healthcheck is interval: 5s, timeout: 5s, retries: 3 (docker-compose.yml), so a ~15s blockage can get the container killed mid-request. There is no timeout on the Gemini call itself (the only guards are the upstream HTTP timeouts: backend OCR_TIMEOUT_MS = 120000, backend/src/features/ocr/external/ocr-client.ts:9, and the frontend's 120s at frontend/src/features/vehicles/api/vehicles.api.ts:90 -- equal values, no headroom), and no retry for transient failures.

This project's asset. The defect has a crisp, cheap-to-measure symptom, an in-repo reference implementation of the fix pattern (extract.py:472), and duration instrumentation already flowing to Loki: backend logs msg="Request processed" with a duration field (backend/src/core/plugins/logging.plugin.ts:20-25), so POST /api/vehicles/decode-vin latency is queryable in Grafana (see mvp-diagnostics-and-logging).

First three steps in this repo:

  1. Baseline with predicted numbers first (Part B discipline). Prediction: with K=3 concurrent VIN decodes in flight, OCR /health response time exceeds 5s and the container restarts. Measure on staging: fire concurrent decodes through the API, while a loop times /health from inside the network (docker exec mvp-ocr-staging curl -s -w '%{time_total}\n' -o /dev/null http://localhost:8000/health) and docker events/restart count watches for healthcheck kills.
  2. Fix the blocking call: in ocr/app/routers/decode.py, wrap the engine call -- await asyncio.get_running_loop().run_in_executor(None, _gemini_engine.decode_vin, vin) -- mirroring extract.py:472; repeat for the sync receipt handlers in extract.py. (The google-genai async client client.aio is an alternative -- UNVERIFIED against the installed SDK version; google-genai>=1.0.0 is an unpinned floor in ocr/requirements.txt:25 and an unsupported config field already broke VIN decode on staging once, per mvp-failure-archaeology. Test any SDK-surface change against the actual deployed image before trusting it.)
  3. Add a service-side timeout shorter than the 120s client timeouts (e.g. 90s around the executor future) plus one retry with jitter for transient Gemini errors, then rerun the step-1 measurement unchanged.

You have a result when during the same saturation experiment as the baseline, /health p99 stays under 1s, the container records zero healthcheck restarts, and VIN decode success latency is unchanged within 10%. Falsified if p99 stays >1s after the executor change -- that means the bottleneck is elsewhere (e.g. thread-pool exhaustion or PaddleOCR model loading), which redirects the work rather than ending it.

Problem 5: Multi-user/SaaS readiness study (candidate -- study, not build)

Why the current state fails. The product is single-tenant by declaration, heading toward paying users; nobody knows how far the gap to 2+ real users actually is, so pricing/positioning decisions (see mvp-launch-readiness) rest on an unexamined assumption in both directions.

This project's asset. The hardest prerequisite is already done: issue #206 migrated all user identity to UUIDs referencing user_profiles.id across 17 feature tables (backend/src/core/identity-migration/migrations/001_migrate_user_id_to_uuid.sql). Note precisely: the columns are still NAMED user_id (the migration renames user_profile_id back to user_id at lines 324+ after backfill) but they now hold user_profiles.id UUIDs -- if any doc says the columns are named user_profile_id, the code wins. Auth already resolves each request to that UUID (backend/src/core/plugins/auth.plugin.ts:165) and repositories scope queries by it. Stripe scaffolding exists per-customer already (subscriptions table with stripe_customer_id, made nullable for admin overrides in backend/src/features/subscriptions/migrations/002_nullable_stripe_customer_id.sql), and an admin_users table exists (backend/src/features/admin/migrations/001_create_admin_users.sql).

First three steps in this repo:

  1. Enumerate candidate blockers by inspection, each phrased as a testable claim. Verified starting list (2026-07-07): no Row Level Security anywhere (zero RLS statements in any migration SQL -- isolation is 100% application-layer WHERE clauses); the tier guard fails open on unknown feature keys (backend/src/core/config/feature-tiers.ts:57-63 returns true for unregistered keys); auth is per-route preHandler, not a global hook, so a forgotten preHandler is an open route; Redis cache keys and the fuel-logs cache invalidation are per-user but the known delete-invalidation miss (imperial-only) would leak stale data per user; vehicle limits free=2/pro=5/enterprise=unlimited (feature-tiers.ts:92-95) assume per-user enforcement paths all consult the same counter.
  2. Run the two-user experiment on staging (never dev/prod without a fresh backup): create a second Auth0 user, then for each feature capsule attempt cross-user reads and writes through the API (user B requests user A's vehicle/fuel-log/document IDs). Record every endpoint that returns another user's data or mutates it. This is the ground truth the study stands on.
  3. Write the go/no-go document as a table: blocker, evidence (the failing/passing test or request transcript from step 2), remediation class (WHERE-clause audit vs RLS vs per-tenant DB vs admin model vs Stripe webhook multiplexing), rough size. File one issue per confirmed blocker, labeled type/chore or type/feature, all status/backlog.

You have a result when the go/no-go document exists with every enumerated blocker backed by a concrete test or transcript -- including the negative results ("tried cross-user read on all 21 capsules' endpoints; N leaked, M correctly 403/404"). A study that finds zero blockers is a result; a study with unproven blockers is not.


PART B -- Methodology: from hunch to accepted result

The evidence bar

A claim (root cause, improvement, feasibility verdict) is accepted here only when:

  1. One mechanism explains ALL observations, including the negatives. If your explanation covers the failure but not why the same code path works elsewhere, it is incomplete. (Full protocol and worked examples: mvp-proof-and-analysis-toolkit.)
  2. It survives assigned adversarial refutation. Before acting, spend one explicit pass (or a spawned subagent) trying to break the conclusion: what observation would disprove it, and did you look for that observation? A conclusion nobody tried to kill is a guess with confidence.
  3. The result is reproducible from the artifacts in the repo -- fixture set, script, LogQL query, or test -- not from a session transcript.

Hypothesis predicts numbers BEFORE running

Write the predicted number down (in the issue or PR description) before running the measurement. "Wrapping the Gemini call in an executor will take /health p99 from >5s to <1s under 3 concurrent decodes" is a hypothesis; "made it async, seems snappier" is not. If you cannot predict a number, you do not yet understand the mechanism -- go back to step 1. After measuring, report predicted vs observed, especially when wrong: a wrong prediction that gets explained is worth more than a right one that does not.

The idea lifecycle in this repo

hunch --> issue --> gated experiment --> adopted | retired
Stage Concrete form here
Hunch A suspicion from a bug, a log pattern, or this file. Costs nothing; commits nothing.
Issue File it via Gitea MCP tools, labeled type/feature or type/chore, status/backlog. State the falsifiable milestone ("result when...") in the body. No sprints/milestones -- work flows directly from issues (owner directive 2026-05-12).
Gated experiment A branch (issue-{n}-{slug}) or a config flag, with the measurable gate declared up front and the predicted numbers written down. Staging is the lab; respect both non-negotiables.
Adopted Merged per mvp-change-control, documented per mvp-docs-and-writing, and -- if it settled a question -- an entry in mvp-failure-archaeology so it is not relitigated.
Retired Written down in mvp-failure-archaeology with the evidence that killed it. A retired idea with documented cause is a settled battle; an undocumented one will be re-fought by a future session that has no memory of yours.

Where good ideas have historically come from in this repo

Use these patterns as prospecting ground, verified in git history:

  • User-pain bug clusters -> systemic fixes. One numeric-as-string bug report (#239, maintenance cost blank on the vehicle summary) triggered an audit that fixed the same defect class in maintenance/ownership-costs (#241) and fuel-logs (#244) within a day, and confirmed stations (#243) already coerced (see mvp-failure-archaeology incident 12) -- and ultimately the detector script and Problem 2 above. When you fix a bug, always ask: where else does this exact shape exist? The audit is usually cheaper than the next bug.
  • Audits -> capsules. The December 2025 security audit (docs/AUDIT.md, now partially stale -- treat its findings as historical, code wins) recommended audit logging; a full backend/src/features/audit-log/ capsule now exists. A written audit with concrete findings reliably converts into shipped features.
  • Incidents -> settled constraints. The google-genai unsupported-config breakage (AFC parameter, reverted in commits 1add6c8/56df5d4) produced the standing rule "test SDK-surface changes against the deployed image" that Problem 4 inherits. Incidents are expensive; extracting a constraint from each one is how the cost is amortized.

Provenance and maintenance

Authored 2026-07-07 from direct repo inspection at commit e729d42 (main). Everything in Part A is a point-in-time snapshot; re-verify before acting. Volatile facts and their re-verification commands:

Fact (as of 2026-07-07) Re-verify with
Issue tracker has zero open issues Gitea MCP list_repo_issues with state=open
CI runs zero tests/lint; PR builds push :latest grep -n "npm test|pytest|latest" .gitea/workflows/staging.yaml
Prod deploy defaults to latest grep -n "default" .gitea/workflows/production.yaml | head -5
Only DATE (OID 1082) parser overridden, not NUMERIC (1700) grep -n setTypeParser backend/src/core/config/database.ts
Numeric-coercion detector exists and its current verdict bash .claude/skills/mvp-diagnostics-and-logging/scripts/check-numeric-coercion.sh
No golden fixtures dir in OCR tests ls ocr/tests/fixtures 2>/dev/null
VIN decode is sync-in-async; executor used only in manual path grep -rn run_in_executor ocr/app and read ocr/app/routers/decode.py
OCR healthcheck 5s/5s/3 grep -A4 healthcheck docker-compose.yml (mvp-ocr block)
google-genai>=1.0.0 unpinned; deployed model is gemini-3-flash-preview grep genai ocr/requirements.txt; grep GEMINI_MODEL docker-compose.yml
Client timeouts both 120s (frontend + backend OCR client) grep -n timeout frontend/src/features/vehicles/api/vehicles.api.ts backend/src/features/ocr/external/ocr-client.ts
User-scoping columns named user_id but hold user_profiles.id UUIDs grep -n "RENAME COLUMN" backend/src/core/identity-migration/migrations/001_migrate_user_id_to_uuid.sql
No RLS anywhere grep -rin "row level security" backend/src --include='*.sql'
Tier guard fails open on unknown keys; limits free=2/pro=5 grep -n -A5 canAccessFeature backend/src/core/config/feature-tiers.ts
stripe_customer_id nullable (admin overrides) ls backend/src/features/subscriptions/migrations/