Removes the old planner/decision-critic/role-agents/domain-agents system (step-injector scripts, stale scopes, sprint-era workflow) and replaces it with 16 ground-truth-verified skills under .claude/skills/: change control, debugging playbook, failure archaeology, architecture contract, domain reference, OCR/Gemini pipeline, config and secrets, build and env, run and operate, diagnostics (with tested scripts), validation and QA, docs and writing, launch readiness, deploy-safety campaign, proof and analysis toolkit, and research frontier. RULE 0/1/2, the temporal-contamination rule, and the decision stress-test protocol are carried forward into the new skills; the retired content remains in git history. Co-Authored-By: Claude Fable 5 <[email protected]>
24 KiB
name, description
| name | description |
|---|---|
| mvp-research-frontier | Open problems where MotoVaultPro can genuinely advance, plus the methodology for turning a hunch into an accepted result. Load when the issue tracker is empty and you need to seed the next work; when someone asks "what should we work on next", "what are the biggest open problems", "is this idea worth pursuing", or "how do I propose/validate a new improvement"; or when starting an experiment, evaluation harness, accuracy measurement, performance study, or multi-tenant/SaaS feasibility study. Trigger keywords - no open issues, roadmap, next steps, research, experiment, hypothesis, golden set, eval harness, accuracy of OCR/VIN extraction, multi-user readiness. |
MotoVaultPro Research Frontier
Everything in this file is labeled candidate/open. Nothing here is a commitment, a roadmap, or owner-approved scope. This skill exists because the issue tracker was fully drained (zero open issues as of 2026-07-07) and the next work must be seeded deliberately, not by whichever bug happens next.
When to use / when NOT to use
Use this skill when:
- You need to propose new work and there is no open issue telling you what to do.
- You are about to start an experiment, measurement, or feasibility study and need the discipline that makes its result acceptable (Part B).
- Someone asks whether one of the five frontier problems below is real or already solved.
Do NOT use this skill for:
- Executing the issue/branch/PR workflow itself --
mvp-change-control. - Diagnosing a live failure --
mvp-debugging-playbook, thenmvp-failure-archaeologyto check whether the battle is already settled. - The CI/deploy-safety hole specifically --
mvp-deploy-safety-campaignis the executable campaign; Problem 3 below is only its extended end-state. - Root-cause evidence standards and benchmark discipline in depth --
mvp-proof-and-analysis-toolkit(this skill cross-references it, not replaces it). - What must be true before launch or public claims --
mvp-launch-readiness. - OCR/Gemini subsystem mechanics --
mvp-ocr-gemini-pipeline. - Running tests / what counts as validation evidence --
mvp-validation-and-qa.
Project reality this skill assumes (owner-confirmed 2026-07-07): development is done by
AI sessions in this repo, the human owner reviews PRs, and end-to-end verification
happens on staging via the PR deploy pipeline because there is no fully working local
dev loop. Two owner non-negotiables apply to every experiment below: (1) no destructive
database operations without a fresh backup (this includes the backend integration tests,
which DROP TABLE ... CASCADE on the shared dev database -- see
backend/src/features/vehicles/tests/integration/vehicles.integration.test.ts:38);
(2) never hand-edit files on staging/prod servers -- rsync --delete on every deploy
reverts them (.gitea/workflows/staging.yaml:117-118, production.yaml:108-109).
PART A -- Frontier problems (all candidate, verified against the repo 2026-07-07)
| # | Problem | One-line asset | Falsifiable milestone |
|---|---|---|---|
| 1 | LLM-extraction evaluation harness | Real extraction pipeline + real receipts | A model-swap PR shows a scorecard diff, not vibes |
| 2 | Type-safe DB boundary | Recurring-bug corpus (#239/#241/#244; #243 reported but not confirmed) + shipped detector script | Detector returns zero suspects AND a regression gate fails a planted bug |
| 3 | CI as a product asset | Working per-PR staging deploy pipeline to extend | A red PR cannot merge; staging never clobbered by unmerged PRs |
| 4 | Async OCR service | Known sync-in-async defect with a measurable symptom | /health p99 < 1s during saturated VIN decode, zero healthcheck restarts |
| 5 | Multi-user/SaaS readiness study | Post-#206 UUID identity scoping already done | Written go/no-go with each blocker proven by a test |
Problem 1: LLM-extraction evaluation harness (candidate)
Why the current state fails. Receipt and VIN extraction accuracy has never been
measured. All OCR tests mock Gemini and generate synthetic images inline with PIL
(e.g. ocr/tests/test_receipt_extraction.py:269-291 builds a blank 100x100 image);
there is no fixtures directory of real labeled images anywhere under ocr/tests/.
Consequences already observed: the deployed model was changed to a preview model
(GEMINI_MODEL: gemini-3-flash-preview, docker-compose.yml:210, overriding the
gemini-2.5-flash code default in ocr/app/config.py) with no before/after accuracy
evidence. Every future model or prompt change is currently judged by eyeball.
This project's asset. A complete, deployed extraction pipeline
(ocr/app/extractors/, ocr/app/routers/{extract,decode}.py) with known output
schemas to score against: fuel receipts produce totalAmount, fuelQuantity,
pricePerUnit, fuelGrade (ocr/app/extractors/fuel_receipt.py:111-147);
maintenance receipts produce serviceName, serviceDate, totalCost, shopName,
laborCost, partsCost, odometerReading
(ocr/app/extractors/maintenance_receipt_extractor.py:40-46); VIN decode produces
year/make/model/trimLevel/engine/transmission (ocr/app/routers/decode.py, response
model in ocr/app/models/schemas.py). Real usage exists: the owner's own uploaded
documents persist in the data/documents/ volume (treat as PII -- get owner sign-off
and strip identifying content before committing any image to git).
First three steps in this repo:
- Create
ocr/tests/fixtures/golden/withreceipts/,maintenance/,vins/subdirectories and onelabels.jsonper subdirectory mapping filename to expected field values (start with N=10 per category; owner-supplied, sanitized images). - Write
ocr/tests/test_golden_extraction.py: for each labeled image, call the extractor and score field-level precision/recall (a field counts as correct only on exact match after normalization; report per-field and aggregate). Gate the whole module behind an env var (e.g.GOLDEN_EVAL=1+pytest.mark.skipif) so the defaultcd ocr && python -m pytestrun stays hermetic -- live Gemini credentials (the WIF chain, seemvp-ocr-gemini-pipeline) exist only in containers. - Run the baseline on staging (the only environment with working credentials): shell into the OCR container and run the gated suite; commit the resulting scorecard JSON next to the fixtures with the model name and date. This number is the baseline every future prompt/model PR must diff against.
You have a result when a PR that changes GEMINI_MODEL, a prompt, or an extractor
includes a before/after scorecard produced by the harness on the same golden set --
and a reviewer can reject it because a number went down. Falsified if, after building
it, scores are so noisy across identical runs (LLM nondeterminism) that diffs are
meaningless; that outcome itself is a publishable finding (then measure variance first
and report mean-of-3).
Problem 2: Type-safe DB boundary -- eliminate the pg-numeric bug class (candidate)
Why the current state fails. node-postgres returns NUMERIC/DECIMAL (PostgreSQL
type OID 1700) as strings. backend/src/core/config/database.ts:12 overrides only the
DATE parser (OID 1082 -> plain string), so every repository mapper must hand-coerce
numerics. This produced at least three shipped bugs (#239, #241, #244 -- #243 was
reported for stations but the stations mappers already coerced via parseFloat, see
mvp-failure-archaeology incident 12) with inconsistent coercion styles across
repositories, and #244 showed a second failure
shape: "enhanced" methods returning raw rows that bypass the mapper entirely. The
convention is enforced by review only; nothing stops the next repository from
reintroducing it.
This project's asset. The bug class is fully characterized (see
mvp-failure-archaeology), and a working detector already ships:
.claude/skills/mvp-diagnostics-and-logging/scripts/check-numeric-coercion.sh
(harvests numeric columns from migration SQL, flags uncoerced row.<col> reads and
raw-row returns in *.repository.ts; exit 0 = clean). There is also an unused codegen
seed: cd backend && npm run schema:generate (backend/src/_system/schema/generate.ts).
First three steps in this repo:
- Run the detector and record the current suspect list as the baseline:
bash .claude/skills/mvp-diagnostics-and-logging/scripts/check-numeric-coercion.sh. - Decide the mechanism with an explicit trade-off writeup (stress-test per
mvp-proof-and-analysis-toolkit). Three options, in ascending ambition: (a) sharedmapNumeric()helper inbackend/src/core/that every mapper calls; (b) global OID 1700 parser inbackend/src/core/config/database.ts-- the OID 1082 override at line 12 is the successful precedent, but weigh honestly: a global parser changes every consumer at once, andparseFloaton money columns trades string safety for float representation (audit currency columns before choosing); (c) schema-derived row types viaschema:generateso uncoerced reads become compile errors. Record the decision and the rejected options in the issue. - Fix all current suspects in one PR, then add the regression gate: wire the detector
into
backend/package.jsonlint (or an eslintno-restricted-syntaxrule inbackend/eslint.config.jsbanningreturn result.rowsin repository files -- candidate, feasibility unproven) so a new violation fails locally and, once Problem 3 lands, in CI.
You have a result when (a) the detector script exits 0 on main, AND (b) a
deliberately planted violation -- an uncoerced row.price read added on a scratch
branch -- is caught by the automated gate without human review. Falsified if the global
parser (option b) breaks an existing consumer that depended on string numerics: that is
a finding, not a failure; document it in mvp-failure-archaeology and fall back to (a).
Problem 3: CI as a product asset (candidate -- end-state extension of the campaign)
Prerequisite: read and follow mvp-deploy-safety-campaign. This entry only records
the frontier beyond that campaign's scope. Do not start here first.
Why the current state fails. CI gates nothing beyond build + boot (canonical
statement: mvp-validation-and-qa section 1). Verified sharp edges: staging.yaml triggers on pull_request (line 12) and every PR
build tags and pushes :latest (lines 58-98), so an unmerged PR overwrites both the
shared staging environment and the very tag that production.yaml deploys by default
(image_tag default 'latest', production.yaml:17). The mobile+desktop hard
requirement is validated nowhere: the single Cypress spec
(frontend/cypress/e2e/stations.cy.ts) has no config and cannot run.
This project's asset. A working self-hosted pipeline that already builds, deploys, and health-verifies a full stack per PR -- most projects have to build that from scratch; here only the gates and isolation are missing.
First three steps in this repo (after the campaign's own steps):
- Split image tagging in
.gitea/workflows/staging.yaml: PR builds push only:{shortsha}(and optionally:pr-{n});:latestmoves only on push tomain. This alone ends the unmerged-code-to-prod path. - Add a test job ahead of
build: backend unit tests (cd backend && npm test -- --forceExit, excludingtests/integrationunless the job provisions an ephemeral postgres+redis -- never point CI tests at a shared database; they DROP TABLE CASCADE), backend/frontendnpm run lintandnpm run type-check, andcd ocr && python -m pytest. Red job blocks merge. - Add a post-deploy viewport smoke job: Playwright (dep already in root
package.json, currently dead -- no config exists) hittinghttps://staging.motovaultpro.comat 375x812 and 1280x800, asserting login page renders and no console errors. This is the first automated teeth behind the mobile+desktop rule.
You have a result when (a) a PR containing a deliberately failing unit test cannot
be merged (workflow red, branch protection enforcing it), AND (b) after a PR build
completes, docker manifest inspect (or registry UI) shows :latest unchanged --
staging mainline is never clobbered by an unmerged PR. Per-PR ephemeral environments or
a namespaced staging are the stretch goal, candidate only; the milestone above does not
require them.
Problem 4: Async OCR service -- Gemini off the event loop (candidate)
Why the current state fails. ocr/app/routers/decode.py calls the synchronous,
unbounded _gemini_engine.decode_vin(vin) directly inside an async def handler
(same pattern for sync receipt paths in routers/extract.py); only the manual-PDF path
uses run_in_executor (extract.py:472, jobs.py:131). A Search-grounded VIN decode
has been observed taking 34s+ (issue #230 history), during which the entire uvicorn
event loop -- including GET /health (ocr/app/main.py:44) -- is blocked. The Docker
healthcheck is interval: 5s, timeout: 5s, retries: 3 (docker-compose.yml), so a
~15s blockage can get the container killed mid-request. There is no timeout on the
Gemini call itself (the only guards are the upstream HTTP timeouts: backend
OCR_TIMEOUT_MS = 120000, backend/src/features/ocr/external/ocr-client.ts:9, and the
frontend's 120s at frontend/src/features/vehicles/api/vehicles.api.ts:90 -- equal
values, no headroom), and no retry for transient failures.
This project's asset. The defect has a crisp, cheap-to-measure symptom, an in-repo
reference implementation of the fix pattern (extract.py:472), and duration
instrumentation already flowing to Loki: backend logs msg="Request processed" with a
duration field (backend/src/core/plugins/logging.plugin.ts:20-25), so
POST /api/vehicles/decode-vin latency is queryable in Grafana (see
mvp-diagnostics-and-logging).
First three steps in this repo:
- Baseline with predicted numbers first (Part B discipline). Prediction: with K=3
concurrent VIN decodes in flight, OCR
/healthresponse time exceeds 5s and the container restarts. Measure on staging: fire concurrent decodes through the API, while a loop times/healthfrom inside the network (docker exec mvp-ocr-staging curl -s -w '%{time_total}\n' -o /dev/null http://localhost:8000/health) anddocker events/restart count watches for healthcheck kills. - Fix the blocking call: in
ocr/app/routers/decode.py, wrap the engine call --await asyncio.get_running_loop().run_in_executor(None, _gemini_engine.decode_vin, vin)-- mirroringextract.py:472; repeat for the sync receipt handlers inextract.py. (The google-genai async clientclient.aiois an alternative -- UNVERIFIED against the installed SDK version;google-genai>=1.0.0is an unpinned floor inocr/requirements.txt:25and an unsupported config field already broke VIN decode on staging once, permvp-failure-archaeology. Test any SDK-surface change against the actual deployed image before trusting it.) - Add a service-side timeout shorter than the 120s client timeouts (e.g. 90s around the executor future) plus one retry with jitter for transient Gemini errors, then rerun the step-1 measurement unchanged.
You have a result when during the same saturation experiment as the baseline,
/health p99 stays under 1s, the container records zero healthcheck restarts, and VIN
decode success latency is unchanged within 10%. Falsified if p99 stays >1s after the
executor change -- that means the bottleneck is elsewhere (e.g. thread-pool exhaustion
or PaddleOCR model loading), which redirects the work rather than ending it.
Problem 5: Multi-user/SaaS readiness study (candidate -- study, not build)
Why the current state fails. The product is single-tenant by declaration, heading
toward paying users; nobody knows how far the gap to 2+ real users actually is, so
pricing/positioning decisions (see mvp-launch-readiness) rest on an unexamined
assumption in both directions.
This project's asset. The hardest prerequisite is already done: issue #206 migrated
all user identity to UUIDs referencing user_profiles.id across 17 feature tables
(backend/src/core/identity-migration/migrations/001_migrate_user_id_to_uuid.sql).
Note precisely: the columns are still NAMED user_id (the migration renames
user_profile_id back to user_id at lines 324+ after backfill) but they now hold
user_profiles.id UUIDs -- if any doc says the columns are named user_profile_id,
the code wins. Auth already resolves each request to that UUID
(backend/src/core/plugins/auth.plugin.ts:165) and repositories scope queries by it.
Stripe scaffolding exists per-customer already (subscriptions table with
stripe_customer_id, made nullable for admin overrides in
backend/src/features/subscriptions/migrations/002_nullable_stripe_customer_id.sql),
and an admin_users table exists (backend/src/features/admin/migrations/001_create_admin_users.sql).
First three steps in this repo:
- Enumerate candidate blockers by inspection, each phrased as a testable claim.
Verified starting list (2026-07-07): no Row Level Security anywhere (zero RLS
statements in any migration SQL -- isolation is 100% application-layer WHERE
clauses); the tier guard fails open on unknown feature keys
(
backend/src/core/config/feature-tiers.ts:57-63returns true for unregistered keys); auth is per-routepreHandler, not a global hook, so a forgotten preHandler is an open route; Redis cache keys and the fuel-logs cache invalidation are per-user but the known delete-invalidation miss (imperial-only) would leak stale data per user; vehicle limits free=2/pro=5/enterprise=unlimited (feature-tiers.ts:92-95) assume per-user enforcement paths all consult the same counter. - Run the two-user experiment on staging (never dev/prod without a fresh backup): create a second Auth0 user, then for each feature capsule attempt cross-user reads and writes through the API (user B requests user A's vehicle/fuel-log/document IDs). Record every endpoint that returns another user's data or mutates it. This is the ground truth the study stands on.
- Write the go/no-go document as a table: blocker, evidence (the failing/passing test
or request transcript from step 2), remediation class (WHERE-clause audit vs RLS vs
per-tenant DB vs admin model vs Stripe webhook multiplexing), rough size. File one
issue per confirmed blocker, labeled
type/choreortype/feature, allstatus/backlog.
You have a result when the go/no-go document exists with every enumerated blocker backed by a concrete test or transcript -- including the negative results ("tried cross-user read on all 21 capsules' endpoints; N leaked, M correctly 403/404"). A study that finds zero blockers is a result; a study with unproven blockers is not.
PART B -- Methodology: from hunch to accepted result
The evidence bar
A claim (root cause, improvement, feasibility verdict) is accepted here only when:
- One mechanism explains ALL observations, including the negatives. If your
explanation covers the failure but not why the same code path works elsewhere, it
is incomplete. (Full protocol and worked examples:
mvp-proof-and-analysis-toolkit.) - It survives assigned adversarial refutation. Before acting, spend one explicit pass (or a spawned subagent) trying to break the conclusion: what observation would disprove it, and did you look for that observation? A conclusion nobody tried to kill is a guess with confidence.
- The result is reproducible from the artifacts in the repo -- fixture set, script, LogQL query, or test -- not from a session transcript.
Hypothesis predicts numbers BEFORE running
Write the predicted number down (in the issue or PR description) before running the measurement. "Wrapping the Gemini call in an executor will take /health p99 from >5s to <1s under 3 concurrent decodes" is a hypothesis; "made it async, seems snappier" is not. If you cannot predict a number, you do not yet understand the mechanism -- go back to step 1. After measuring, report predicted vs observed, especially when wrong: a wrong prediction that gets explained is worth more than a right one that does not.
The idea lifecycle in this repo
hunch --> issue --> gated experiment --> adopted | retired
| Stage | Concrete form here |
|---|---|
| Hunch | A suspicion from a bug, a log pattern, or this file. Costs nothing; commits nothing. |
| Issue | File it via Gitea MCP tools, labeled type/feature or type/chore, status/backlog. State the falsifiable milestone ("result when...") in the body. No sprints/milestones -- work flows directly from issues (owner directive 2026-05-12). |
| Gated experiment | A branch (issue-{n}-{slug}) or a config flag, with the measurable gate declared up front and the predicted numbers written down. Staging is the lab; respect both non-negotiables. |
| Adopted | Merged per mvp-change-control, documented per mvp-docs-and-writing, and -- if it settled a question -- an entry in mvp-failure-archaeology so it is not relitigated. |
| Retired | Written down in mvp-failure-archaeology with the evidence that killed it. A retired idea with documented cause is a settled battle; an undocumented one will be re-fought by a future session that has no memory of yours. |
Where good ideas have historically come from in this repo
Use these patterns as prospecting ground, verified in git history:
- User-pain bug clusters -> systemic fixes. One numeric-as-string bug report
(#239, maintenance cost blank on the vehicle summary) triggered an audit that fixed
the same defect class in maintenance/ownership-costs (#241) and fuel-logs (#244)
within a day, and confirmed stations (#243) already coerced (see
mvp-failure-archaeologyincident 12) -- and ultimately the detector script and Problem 2 above. When you fix a bug, always ask: where else does this exact shape exist? The audit is usually cheaper than the next bug. - Audits -> capsules. The December 2025 security audit (
docs/AUDIT.md, now partially stale -- treat its findings as historical, code wins) recommended audit logging; a fullbackend/src/features/audit-log/capsule now exists. A written audit with concrete findings reliably converts into shipped features. - Incidents -> settled constraints. The google-genai unsupported-config breakage
(AFC parameter, reverted in commits
1add6c8/56df5d4) produced the standing rule "test SDK-surface changes against the deployed image" that Problem 4 inherits. Incidents are expensive; extracting a constraint from each one is how the cost is amortized.
Provenance and maintenance
Authored 2026-07-07 from direct repo inspection at commit e729d42 (main). Everything
in Part A is a point-in-time snapshot; re-verify before acting. Volatile facts and
their re-verification commands:
| Fact (as of 2026-07-07) | Re-verify with |
|---|---|
| Issue tracker has zero open issues | Gitea MCP list_repo_issues with state=open |
CI runs zero tests/lint; PR builds push :latest |
grep -n "npm test|pytest|latest" .gitea/workflows/staging.yaml |
Prod deploy defaults to latest |
grep -n "default" .gitea/workflows/production.yaml | head -5 |
| Only DATE (OID 1082) parser overridden, not NUMERIC (1700) | grep -n setTypeParser backend/src/core/config/database.ts |
| Numeric-coercion detector exists and its current verdict | bash .claude/skills/mvp-diagnostics-and-logging/scripts/check-numeric-coercion.sh |
| No golden fixtures dir in OCR tests | ls ocr/tests/fixtures 2>/dev/null |
| VIN decode is sync-in-async; executor used only in manual path | grep -rn run_in_executor ocr/app and read ocr/app/routers/decode.py |
| OCR healthcheck 5s/5s/3 | grep -A4 healthcheck docker-compose.yml (mvp-ocr block) |
google-genai>=1.0.0 unpinned; deployed model is gemini-3-flash-preview |
grep genai ocr/requirements.txt; grep GEMINI_MODEL docker-compose.yml |
| Client timeouts both 120s (frontend + backend OCR client) | grep -n timeout frontend/src/features/vehicles/api/vehicles.api.ts backend/src/features/ocr/external/ocr-client.ts |
User-scoping columns named user_id but hold user_profiles.id UUIDs |
grep -n "RENAME COLUMN" backend/src/core/identity-migration/migrations/001_migrate_user_id_to_uuid.sql |
| No RLS anywhere | grep -rin "row level security" backend/src --include='*.sql' |
| Tier guard fails open on unknown keys; limits free=2/pro=5 | grep -n -A5 canAccessFeature backend/src/core/config/feature-tiers.ts |
stripe_customer_id nullable (admin overrides) |
ls backend/src/features/subscriptions/migrations/ |