TF AI-QC — Complete Build Guide (Google Cloud + Vertex AI)
Version: 2.0 — 2026-06-28 | Replaces: GOOGLE-CLOUD-GUIDE.md (v1.0)
Stack decision: Google Cloud Platform, all services. Railway/Cloudflare stack is deprecated.
What Changed From v1.0
| Item | v1.0 | v2.0 |
|---|---|---|
| Vertex AI model | Gemini 1.5 Pro | Gemini 2.5 Pro (use gemini-2.5-pro) |
| Auth | Custom JWT | Firebase Auth + Google OAuth (Google Workspace SSO) |
| Stack decision | Two options documented | GCP only — decision made |
| Context folder | Listed as needed | Explicit loading steps added (Step 0) |
| PII scrubber | Mentioned | Full implementation included |
| Bubble integration | Session 8, empty folder | Still Session 8 — get API docs first |
Stack Summary
| Component | Service | Why |
|---|---|---|
| Backend | Cloud Run (FastAPI/Python) | Serverless, HTTPS automatic, scales to zero |
| Database | Cloud SQL PostgreSQL 15 | Managed, SSL required, encryption at rest by default |
| File storage | Google Cloud Storage | AES-256 at rest, signed URLs, GLBA-ready |
| Auth | Firebase Authentication + Google OAuth | SSO with @truefootage.com accounts, no passwords |
| Frontend | Firebase Hosting (React/TypeScript/Tailwind) | Free, CDN-backed, integrates with Cloud Run backend |
| AI scoring | Vertex AI — Gemini 2.5 Pro | Data stays in Google's boundary, DPA covered by GCP ToS |
| Background jobs | Cloud Tasks | QC jobs queued and retried automatically |
| Secrets | Secret Manager | No .env files in production |
| Logs / audit | Cloud Logging | Append-only, GLBA audit trail |
| CI/CD | Cloud Build | Auto-deploy on push to main |
| SendGrid (free tier) | 3,000 emails/mo free, sufficient for beta |
GLBA compliance: Google's DPA (accepted when you accept GCP Terms of Service) covers all services above. No separate DPA negotiation required for beta. Before production, explicitly accept at: GCP Console → IAM → Privacy.
MASTER PROGRESS CHECKLIST
PHASE 0: Pre-Session Setup (Do First — Sessions Will Fail Without This)
- ☐ P0. Load UAD 3.6 Supplement PDF into
context/rule-references/uad-3.6/ - ☐ P1. Load Fannie Mae Selling Guide B4-1 PDF into
context/guidelines/gse/ - ☐ P2. Load FHA 4000.1 Chapter II.D PDF into
context/guidelines/gse/ - ☐ P3. Load VA Lender Handbook Chapter 11 PDF into
context/guidelines/gse/ - ☐ P4. Load USPAP current edition PDF into
context/guidelines/uspap/ - ☐ P5. Add redacted URAR sample (XML preferred, PDF accepted) to
context/sample-reports/1004/ - ☐ P6. Strip ALL PII from sample report before adding (remove borrower name, SSN, address, loan number)
PHASE A: Google Cloud Account Setup
- ☐ A1. Create Google Cloud account at cloud.google.com
- ☐ A2. Create GCP project: ID =
tf-ai-qc-prod - ☐ A3. Enable billing, set alert at $200/month
- ☐ A4. Enable all required APIs (command in Section 1)
- ☐ A5. Install gcloud CLI
- ☐ A6. Install Firebase CLI (
npm install -g firebase-tools) - ☐ A7. Authenticate both CLIs with your Google account
- ☐ A8. Set default project:
gcloud config set project tf-ai-qc-prod
PHASE B: Core Services
- ☐ B1. Create Cloud SQL instance:
tfaiqc-db, PostgreSQL 15,us-central1 - ☐ B2. Create database:
tfaiqc - ☐ B3. Save Cloud SQL credentials (instance, db name, user, password, connection name)
- ☐ B4. Create GCS bucket:
tf-ai-qc-reports-prod— enforce public access prevention - ☐ B5. Enable Firebase on
tf-ai-qc-prodproject - ☐ B6. Enable Firebase Auth with Google provider (restrict to
@truefootage.com) - ☐ B7. Create Firebase Hosting site
- ☐ B8. Verify Vertex AI API enabled, confirm Gemini 2.5 Pro in Model Garden
- ☐ B9. Create Secret Manager secrets (list in Section 2)
- ☐ B10. Create service account
tfaiqc-backendwith minimum roles (list in Section 2) - ☐ B11. Create Artifact Registry repo:
tfaiqc-images, Docker,us-central1
PHASE C: Development Environment
- ☐ C1. Python 3.12 installed (
python --version) - ☐ C2. Node.js 18+ installed (
node --version) - ☐ C3. VS Code with Python, Pylance, Tailwind CSS IntelliSense, GitLens, REST Client extensions
- ☐ C4. Open
C:\Users\kzele\Claude Cowork\Projects\TF AI-QCin VS Code - ☐ C5. Create GitHub repo:
TF-AI-QC-GCP(separate from original TF-AI-QC repo) - ☐ C6. Push project folder to new repo
PHASE D: Build Sessions (20 sessions)
- ☐ D0. Session 0 — GCP Project scaffold
- ☐ D1a. Session 1A — Database models
- ☐ D1b. Session 1B — Firebase Authentication
- ☐ D1c. Session 1C — Cloud Storage file upload
- ☐ D2a. Session 2A — XML parser (UAD 3.6)
- ☐ D2b. Session 2B — PDF extractor
- ☐ D3a. Session 3A — Rule engine framework + UAD formatting rules
- ☐ D3b. Session 3B — GSE overlay rules
- ☐ D3c. Session 3C — Quality scoring + Vertex AI (Gemini 2.5 Pro)
- ☐ D3d. Session 3D — Rule admin + full pipeline wiring
- ☐ D4a. Session 4A — Workflow state machine
- ☐ D4b. Session 4B — Revision request system
- ☐ D5a. Session 5A — Appraiser portal (frontend)
- ☐ D5b. Session 5B — Reviewer dashboard (frontend)
- ☐ D5c. Session 5C — Admin panel (frontend)
- ☐ D5d. Session 5D — Coaching dashboards (frontend)
- ☐ D6. Session 6 — Deploy to Cloud Run + Firebase Hosting
- ☐ D7a. Session 7A — Pattern detection engine
- ☐ D7b. Session 7B — Coaching reports + export
- ☐ D8. Session 8 — Bubble.io OMS integration (requires Bubble API docs first)
PHASE E: Launch Prep
- ☐ E1. Security checklist review (Section 7)
- ☐ E2. Confirm Google Cloud DPA accepted at GCP Console → IAM → Privacy
- ☐ E3. Beta test with 5 appraisers
- ☐ E4. Verify audit logging active and tested
- ☐ E5. Set billing alert at $150/month
- ☐ E6. Publish Firebase Hosting URL to team
SECTION 0 — Load Context Files (Do Before Session 0)
Sessions 2A, 3A, and 3B reference context/ files. If the folders are empty, those sessions produce generic code instead of UAD 3.6–specific logic.
What to Load and Where
context/
├── rule-references/
│ └── uad-3.6/
│ └── UAD-3.6-Supplement-2026-06-03.pdf ← Fannie Mae/Freddie Mac UAD 3.6 Supplement
├── guidelines/
│ ├── gse/
│ │ ├── FannieMae-SellingGuide-B4-1.pdf ← Fannie Mae Selling Guide, Subpart B4-1
│ │ ├── FHA-4000.1-Chapter-IID.pdf ← FHA Handbook Chapter II.D
│ │ └── VA-LenderHandbook-Chapter11.pdf ← VA Lender Handbook Chapter 11
│ └── uspap/
│ └── USPAP-2024-2025.pdf ← Current USPAP edition
└── sample-reports/
└── 1004/
└── sample-urar-redacted.xml ← Redacted UAD 3.6 URAR (XML preferred)
Where to Get Each Document
| Document | Source |
|---|---|
| UAD 3.6 Supplement | Already in your UAD 3.6 Wiki project folder — copy it here |
| Fannie Mae Selling Guide B4-1 | selling-guide.fanniemae.com → search B4-1 → export PDF |
| FHA 4000.1 Chapter II.D | hud.gov → search "4000.1 handbook" → Chapter II.D |
| VA Lender Handbook Chapter 11 | benefits.va.gov → search "VA Lender Handbook Chapter 11" |
| USPAP | appraisalfoundation.org → Publications → USPAP (login required) |
| Sample URAR XML | Export from TOTAL, ACI, or ClickForms — any completed report. Strip PII. |
PII Stripping for Sample Report
Before adding any sample report to context/:
- Open the XML in Notepad++ or VS Code
- Find and replace:
- Borrower name →
SAMPLE BORROWER - Property address →
123 SAMPLE ST, ANYTOWN, ST 00000 - SSN / loan number →
000000000 - Lender name →
SAMPLE LENDER - Appraiser license number →
SAMPLE-LICENSE-001
- Borrower name →
- Save as
sample-urar-redacted.xml
SECTION 1 — GCP Account Setup
Enable All Required APIs (One Command)
In Cloud Shell (>_ icon in GCP Console) or PowerShell with gcloud installed:
gcloud services enable \
run.googleapis.com \
sqladmin.googleapis.com \
storage.googleapis.com \
aiplatform.googleapis.com \
firebase.googleapis.com \
iam.googleapis.com \
secretmanager.googleapis.com \
cloudbuild.googleapis.com \
artifactregistry.googleapis.com \
cloudtasks.googleapis.com \
logging.googleapis.com \
monitoring.googleapis.com \
--project=tf-ai-qc-prod
Set Default Project
gcloud config set project tf-ai-qc-prod
gcloud config set compute/region us-central1
SECTION 2 — Core Services Setup
Cloud SQL
Instance ID: tfaiqc-db
Database version: PostgreSQL 15
Region: us-central1
Machine type: db-f1-micro (beta) → db-g1-small (production)
Database name: tfaiqc
User: postgres
Connection name: tf-ai-qc-prod:us-central1:tfaiqc-db
Enable SSL connections: Cloud SQL instance → Connections → SSL → Require SSL
Cloud Storage Bucket
Name: tf-ai-qc-reports-prod
Region: us-central1
Storage class: Standard
Access control: Uniform
Public access: BLOCKED (enforce public access prevention ✅)
Encryption: Google-managed (default, AES-256)
Firebase Authentication
- Firebase Console → Authentication → Sign-in method → Google → Enable
- Restrict to domain: In Google sign-in settings → Authorized domains → add
truefootage.com - Also enable Email/Password for non-Google accounts
Secret Manager — Secrets to Create
Secret Name Value
────────────────────── ────────────────────────────────────────────
db-password Cloud SQL postgres password
sendgrid-api-key SendGrid API key
firebase-admin-key Firebase service account JSON (full contents)
secret-key Random 32-byte base64 string (see below)
bubble-api-token Leave empty for now — populate before Session 8
Generate secret-key:
[System.Convert]::ToBase64String([System.Security.Cryptography.RandomNumberGenerator]::GetBytes(32))
Service Account
Name: tfaiqc-backend
Roles to assign:
roles/cloudsql.clientroles/storage.objectAdminroles/aiplatform.userroles/cloudtasks.enqueuerroles/logging.logWriterroles/secretmanager.secretAccessor
After creation: Keys tab → Add Key → JSON → download → paste contents into firebase-admin-key secret → delete downloaded file.
SECTION 3 — Session Rules (Read Before Every Session)
Before Every Session
- Open Claude Code (Cowork)
- Switch model to claude-opus-4-8
- Enable Extended Thinking (High Effort / Extended Thinking toggle ON)
- Select the
TF AI-QCproject folder - Paste the session prompt as your first message
- Do not interrupt mid-session — let it complete the full implementation
After Every Session
- Review output in VS Code
- Run the verification commands Claude provides
- Commit:
git add . git commit -m "Session X: [description]" git push origin main - Check off the session in Phase D checklist
If a Session Times Out or Fails
- Do not start a new session from scratch
- Open a new conversation, paste this header first:
I am continuing TF AI-QC on Google Cloud. The spec is at: docs/superpowers/specs/2026-06-25-uad-qc-tool-design.md GCP Project: tf-ai-qc-prod | Region: us-central1 Previous session ended at: [describe where it stopped] Continue from that point.
SECTION 4 — All Session Prompts (Updated for Vertex AI + Gemini 2.5 Pro)
SESSION 0 — GCP Project Scaffold
You are helping me build TF AI-QC, a residential appraisal QC tool for True Footage, running entirely on Google Cloud Platform. Use extended thinking for this session.
Read the design spec first:
docs/superpowers/specs/2026-06-25-uad-qc-tool-design.md
TECH STACK:
- Backend: Python 3.12 / FastAPI — Google Cloud Run
- Frontend: React / TypeScript / Tailwind — Firebase Hosting
- Database: Google Cloud SQL (PostgreSQL 15)
- File Storage: Google Cloud Storage
- Authentication: Firebase Authentication with Google OAuth
- AI: Google Vertex AI with Gemini 2.5 Pro (model ID: gemini-2.5-pro)
- Secrets: Google Secret Manager
- Background jobs: Google Cloud Tasks
- Monitoring: Google Cloud Logging
GCP Project ID: tf-ai-qc-prod
Region: us-central1
GCS Bucket: tf-ai-qc-reports-prod
Cloud SQL: tfaiqc-db / database: tfaiqc
Artifact Registry: us-central1-docker.pkg.dev/tf-ai-qc-prod/tfaiqc-images/backend
Set up the complete project scaffold:
BACKEND (Python/FastAPI):
- backend/pyproject.toml with all dependencies:
fastapi, uvicorn, sqlalchemy>=2.0, alembic, cloud-sql-python-connector[pg8000],
firebase-admin, google-cloud-storage, google-cloud-aiplatform, google-cloud-tasks,
google-cloud-secret-manager, google-cloud-logging, python-multipart, pydantic-settings,
passlib[bcrypt], lxml, pdfplumber, httpx, pytest, pytest-asyncio
- backend/app/main.py — FastAPI app, CORS restricted to Firebase Hosting domain, routers, lifespan events
- backend/app/core/config.py — Settings using pydantic-settings, loads from Secret Manager in production, .env.local for dev
- backend/app/core/gcp.py — initialize Storage, Tasks, Secret Manager, Logging clients at startup
- backend/app/core/database.py — Cloud SQL connector setup (Unix socket in Cloud Run, TCP for local dev via Cloud SQL Auth Proxy)
- backend/.env.local.example — all required env vars documented
- Dockerfile — multi-stage build, non-root user, port 8080
- cloudbuild.yaml — test → build → push to Artifact Registry → deploy to Cloud Run → migrate DB → deploy frontend
FRONTEND (React/TypeScript):
- frontend/ initialized with Vite + React + TypeScript
- Dependencies: react-router-dom, axios, @tanstack/react-query, zustand, tailwindcss, firebase
- frontend/src/firebase.ts — Firebase app init
- frontend/src/auth/ — AuthProvider, useAuth hook, GoogleSignInButton
- frontend/src/api/client.ts — axios instance that auto-attaches Firebase ID token
- .firebaserc pointing to tf-ai-qc-prod
- firebase.json with hosting config and Cloud Run rewrite for /api/*
DATABASE:
- Alembic configured for Cloud SQL (Unix socket production, TCP local)
- Initial migration: users, reports, qc_results, qc_flags, revisions, revision_responses, rules tables
- Match data model from spec Section 5
Verify backend starts: uvicorn app.main:app --reload
Show output.
SESSION 1A — Database Models
You are building TF AI-QC on Google Cloud. Use extended thinking.
Read: docs/superpowers/specs/2026-06-25-uad-qc-tool-design.md
Scaffold is in place. Create SQLAlchemy 2.x ORM models using DeclarativeBase syntax.
backend/app/models/user.py — User
- id (UUID primary key), email, name
- google_uid (unique — Firebase UID, primary identifier)
- role: Enum("appraiser", "reviewer", "admin")
- license_number, license_state, active (bool, default True)
- created_at, updated_at
backend/app/models/report.py — Report
- id (UUID), uploader_id (FK users), file_type: Enum("xml", "pdf")
- form_type: Enum("1004", "1073", "1025", "2055")
- gcs_path (GCS object path — NOT a signed URL), file_size_bytes
- subject_address, status: Enum("submitted","qc_running","qc_complete","revision_requested","resubmitted","approved")
- bubble_order_id (nullable — for OMS integration later)
- submitted_at, completed_at, run_number (default 1)
backend/app/models/qc_result.py — QCResult
- id (UUID), report_id (FK), run_number
- hard_pass (bool), quality_score (Numeric 0-100)
- raw_flags (JSONB), vertex_ai_score (Numeric, nullable), vertex_ai_flags (JSONB, nullable)
- duration_ms (int), created_at
backend/app/models/qc_flag.py — QCFlag
- id (UUID), qc_result_id (FK)
- category: Enum("compliance", "quality")
- severity: Enum("fail", "warning", "info")
- field_name, rule_code, message, gse_reference
- status: Enum("open", "waived", "resolved"), waiver_reason (nullable), waived_by (FK users, nullable)
backend/app/models/revision.py — Revision + RevisionResponse
- Revision: id, report_id (FK), flag_id (FK nullable), created_by (FK reviewer)
- message, priority: Enum("required", "recommended"), status: Enum("open", "responded", "closed")
- RevisionResponse: id, revision_id (FK), created_by (FK appraiser), message, created_at
backend/app/models/rule.py — Rule
- id (UUID), rule_code (unique), category, severity, description
- gse_applicability (ARRAY of text), active (bool), config (JSONB), created_at
backend/app/models/__init__.py — export all
Run Alembic migration against local dev database.
Show output confirming all tables created.
SESSION 1B — Firebase Authentication
You are building TF AI-QC on Google Cloud. Use extended thinking.
Read: docs/superpowers/specs/2026-06-25-uad-qc-tool-design.md
Database models are done. Build Firebase Authentication integration.
backend/app/core/firebase_auth.py:
- Initialize Firebase Admin SDK using credentials from Secret Manager (key: firebase-admin-key)
- verify_firebase_token(token: str) → decoded token dict or raise HTTPException 401
- Extract: uid (google_uid), email, name from token claims
backend/app/core/dependencies.py:
- get_current_user(token from Authorization: Bearer header) → User DB record
1. Verify token via firebase_auth.verify_firebase_token()
2. Look up User by google_uid
3. If not found and token is valid, create User with role="appraiser" (first login auto-provision)
4. If user.active is False, raise 403
- require_reviewer: dependency that calls get_current_user and checks role in ["reviewer", "admin"]
- require_admin: dependency that checks role == "admin"
backend/app/api/routes/auth.py:
- GET /auth/me → current user profile (uses get_current_user)
- POST /auth/sync → create or update User from Firebase token (called on every app load)
Updates: name, email (in case Google account name changed)
frontend/src/auth/AuthProvider.tsx:
- Firebase Auth context: auth state, token (refreshed every 55 min before 1h expiry), loading
- On token refresh: call POST /auth/sync to keep DB in sync
frontend/src/auth/GoogleSignIn.tsx:
- signInWithPopup with GoogleAuthProvider
- After sign-in: call /auth/sync, redirect to dashboard
frontend/src/auth/useAuth.ts:
- Returns { user, token, isAuthenticated, signIn, signOut, loading }
frontend/src/api/client.ts:
- Axios instance with baseURL = import.meta.env.VITE_API_URL
- Request interceptor: attach Authorization: Bearer {token} to every request
- Response interceptor: on 401, sign out user and redirect to login
Write tests: mock Firebase Admin SDK, test token verification and user auto-provisioning.
SESSION 1C — Google Cloud Storage File Upload
You are building TF AI-QC on Google Cloud. Use extended thinking.
Read: docs/superpowers/specs/2026-06-25-uad-qc-tool-design.md
Auth is working. Build file upload using Google Cloud Storage.
backend/app/services/storage.py — GCSStorageService:
- upload_report(file_bytes, original_filename, content_type, user_id) → gcs_path
Path format: reports/{user_id}/{YYYY-MM}/{uuid}.{ext}
- generate_signed_url(gcs_path, expiry_minutes=15) → signed URL
Use v4 signing. Expiry MUST be 15 minutes — enforced, not configurable.
- delete_file(gcs_path)
- NEVER log file contents. Log only: user_id, gcs_path (no signed URLs in logs), file_size.
backend/app/api/routes/reports.py:
- POST /reports/upload (multipart, requires auth):
1. Validate content-type header: reject if not XML or PDF
2. Validate file size: reject >50MB
3. Validate magic bytes: XML must start with <?xml or <MISMO, PDF must start with %PDF
4. Upload to GCS via GCSStorageService
5. Create Report record: status=submitted, gcs_path stored (not signed URL)
6. Enqueue QC Cloud Task
7. Log: user_id, report_id, file_type, file_size_bytes, timestamp — NOTHING ELSE
8. Return: { report_id, status, form_type_detected }
- GET /reports (paginated, filterable by status):
Appraiser: own reports only (WHERE uploader_id = current_user.id)
Reviewer/Admin: all reports
- GET /reports/{id}:
Generate fresh signed URL (15-min expiry) on each request
Include: report metadata, latest qc_result summary, open revision count
Log: user_id, report_id, "view_report", timestamp
Privacy rule: signed URLs are generated at request time only. gcs_path is never exposed to frontend. Permanent GCS links never returned.
Integration test: upload a 1KB test XML file, confirm GCS object exists and DB record created.
SESSION 2A — UAD 3.6 XML Parser
You are building TF AI-QC on Google Cloud. Use extended thinking for complex parsing logic.
Read ALL of these before writing any code:
- docs/superpowers/specs/2026-06-25-uad-qc-tool-design.md
- context/rule-references/uad-3.6/ (all files — read them)
- context/sample-reports/1004/ (all files — read them)
Build the UAD 3.6 XML ingest engine.
backend/app/services/ingest/models.py — ReportData dataclass:
- subject_property: SubjectProperty (address, legal_desc, property_type, year_built, gla, site_size, condition C1-C6, quality Q1-Q6, age_roof)
- contract: Contract (sale_price, date, financing_type, seller_concessions)
- neighborhood: Neighborhood (name, boundaries, built_up, growth, supply_demand, marketing_time, trend, price_range_low, price_range_high, one_year_change)
- site: Site (dimensions, zoning, utilities, flood_zone, fema_map_number)
- improvements: Improvements (foundation, exterior, roof, interior, heating, cooling, rooms, bedrooms, baths, amenities list)
- comparables: List[Comparable] (up to 6 — address, proximity, sale_price, sale_date, gla, site, age, condition, adjustments dict, adjusted_price, data_source)
- reconciliation: Reconciliation (cost_value, sales_comparison_value, income_value, final_value, value_type, reconciliation_comment)
- certifications: Certifications (statements list, appraiser_name, license_number, license_state, license_type, effective_date, supervisor_name, supervisor_license)
- new_fields: NewUAD36Fields (non_residential_use, water_frontage, disaster_mitigation, energy_features, front_door_elevation, converted_area, noncontinuous_area, kitchen_bath_details, defects_list, outbuilding_details)
- field_confidence: Dict[str, float] (1.0=explicit, 0.5=derived, 0.0=absent)
- low_confidence_fields: List[str]
- raw_fields: Dict[str, Any] (all unmapped fields)
- extraction_warnings: List[str]
backend/app/services/ingest/xml_parser.py — XMLParser:
- parse(xml_bytes: bytes) → ReportData
- Use lxml for namespace-aware parsing
- Handle MISMO UAD 3.6 schema namespaces
- Distinguish empty string from absent field (different QC implications — absent = missing, empty = left blank)
- For each field: confidence 1.0 if explicitly present with value, 0.5 if derived, 0.0 if absent
Unit tests with synthetic UAD 3.6 XML fixture. Run and show output.
SESSION 2B — PDF Extractor
You are building TF AI-QC on Google Cloud. Use extended thinking.
Read ALL of these before writing any code:
- docs/superpowers/specs/2026-06-25-uad-qc-tool-design.md
- context/sample-reports/ (all files in all subdirectories)
XML parser is complete. Build PDF extraction engine for the new URAR form.
backend/app/services/ingest/pdf_extractor.py — PDFExtractor:
- extract(pdf_bytes: bytes) → ReportData
- Use pdfplumber for text and table extraction
- Detect appraisal software by checking PDF metadata creator field and text patterns:
- TOTAL: metadata creator contains "TOTAL"
- ACI: contains "ACI"
- ClickForms: contains "ClickForms"
- Unknown: fall back to generic label-proximity detection
- Apply software-specific coordinate profiles for field location
- Assign confidence per field: 1.0 (clear label-value pair), 0.7 (derived from table), 0.5 (position-inferred), 0.0 (not found)
- Flag fields with confidence < 0.7 as low_confidence
backend/app/services/ingest/extractor_factory.py — ExtractorFactory:
- detect_file_type(file_bytes: bytes) → Literal["xml", "pdf"]
Check magic bytes only — NOT filename or content-type header
XML: starts with b"<?xml" or b"<MISMO" or b"\xef\xbb\xbf<?xml" (UTF-8 BOM)
PDF: starts with b"%PDF"
- extract(file_bytes: bytes) → tuple[ReportData, str] (report_data, detected_type)
Unit tests. If no sample PDF in context/, generate a minimal test PDF using reportlab.
SESSION 3A — Rule Engine Framework + UAD Formatting Rules
You are building TF AI-QC on Google Cloud. Use extended thinking — the rule engine architecture is the core of the product. Think carefully before writing code.
Read ALL of these before writing:
- docs/superpowers/specs/2026-06-25-uad-qc-tool-design.md
- context/rule-references/uad-3.6/ (all files)
Build the rule engine framework and UAD 3.6 formatting rules.
backend/app/services/rules/base_rule.py:
- RuleViolation dataclass: rule_code, field_name, message, gse_reference, severity, evidence (what was found vs. expected)
- BaseRule abstract class:
- Class attrs: rule_code (str), category (Literal["compliance","quality"]), severity (Literal["fail","warning","info"]), description, gse_applicability (list[str])
- Abstract: evaluate(report_data: ReportData) → list[RuleViolation]
backend/app/services/rules/engine.py — RuleEngine:
- __init__: load active rules from DB, cache with 5-min TTL
- run(report_data: ReportData) → RuleEngineResult
Pass 1 (compliance): run all rules where category=="compliance". Collect violations. hard_pass = no severity=="fail" violations.
Pass 2 (quality): run all rules where category=="quality". Produce 0-100 score.
- RuleEngineResult dataclass: hard_pass, quality_score, violations list, duration_ms, rules_run_count
- refresh_cache(): force reload from DB (called by rule admin API)
UAD Formatting Rules — backend/app/services/rules/uad_format/:
rule_date_format.py — All dates must be MM/DD/YYYY
rule_condition_rating.py — Condition must be one of C1, C2, C3, C4, C5, C6 only
rule_quality_rating.py — Quality must be one of Q1, Q2, Q3, Q4, Q5, Q6 only
rule_required_fields_urar.py — All required URAR fields non-empty for UAD 3.6 new form
Required: subject address, legal description, property type, year built, GLA, condition, quality,
neighborhood section complete, min 3 comps, reconciliation non-empty, certification statements
rule_uad_abbreviations.py — Validate UAD valid values:
built_up: Over 75%, 25-75%, Under 25%
growth: Rapid, Stable, Slow
supply_demand: Shortage, In Balance, Over Supply
marketing_time: Under 3 Months, 3-6 Months, Over 6 Months
trend: Increasing, Stable, Declining
rule_value_positive.py — Final appraised value must be positive number
rule_comp_minimum.py — Minimum 3 comparable sales required (UAD 3.6 requirement)
rule_certification_present.py — USPAP certification statement must be non-empty and contain required language
Comprehensive unit tests for every rule with both passing and failing inputs.
Run tests and show pass/fail counts.
SESSION 3B — GSE Overlay Rules
You are building TF AI-QC on Google Cloud. Use extended thinking.
Read ALL of these before writing rules:
- docs/superpowers/specs/2026-06-25-uad-qc-tool-design.md
- context/guidelines/gse/ (all files — read BEFORE writing any rule)
- context/rule-references/uad-3.6/ (all files)
Build GSE-specific rules. Each rule must declare gse_applicability list accurately from the actual guidance.
Fannie Mae — backend/app/services/rules/gse/fannie_mae/:
rule_fnma_comp_count.py — min 3, note <6 for complex properties (gse=["FNMA"])
rule_fnma_comp_age.py — comps >12 months require written explanation in report (gse=["FNMA"])
rule_fnma_comp_proximity.py — >1 mile urban / >5 miles rural requires explanation (gse=["FNMA"])
rule_fnma_market_conditions.py — declining market trend requires additional market support (gse=["FNMA"])
rule_fnma_gla_adjustment.py — GLA difference >15% between comp and subject requires adjustment (gse=["FNMA"])
rule_fnma_net_adjustment.py — net adjustment >15% of comp sale price: warning (gse=["FNMA"])
rule_fnma_gross_adjustment.py — gross adjustment >25%: warning (gse=["FNMA"])
rule_fnma_concession_adjustment.py — seller concessions must be adjusted to reflect market (gse=["FNMA"])
Freddie Mac — gse/freddie_mac/:
rule_fhlmc_comp_selection.py — comps must be most recent and similar available (gse=["FHLMC"])
rule_fhlmc_age_adjustment.py — time adjustments required if market changed direction (gse=["FHLMC"])
FHA — gse/fha/:
rule_fha_condition.py — C5 or C6 condition triggers repair requirement flag (gse=["FHA"])
rule_fha_well_septic.py — well or septic requires specific certification language (gse=["FHA"])
rule_fha_flood.py — AE or VE flood zone requires flood certification note (gse=["FHA"])
VA — gse/va/:
rule_va_mpr.py — VA MPR flags for condition issues (gse=["VA"])
rule_va_tidewater.py — value indication below contract price triggers Tidewater initiation flag (gse=["VA"])
USPAP — backend/app/services/rules/uspap/:
rule_uspap_certification.py — all required USPAP certifications present and non-boilerplate
rule_uspap_scope_of_work.py — scope of work statement non-empty and property-specific
rule_uspap_limiting_conditions.py — limiting conditions statement present
rule_uspap_appraiser_credentials.py — license number, state, type all populated
Run all tests. Show total pass/fail counts by rule.
SESSION 3C — Quality Scoring + Vertex AI (Gemini 2.5 Pro)
You are building TF AI-QC on Google Cloud. Use extended thinking — this is the most sensitive data-handling session. Read:
docs/superpowers/specs/2026-06-25-uad-qc-tool-design.md
Build the quality scoring engine with Vertex AI integration.
PII SCRUBBER (implement this first — required before any text reaches Vertex AI):
backend/app/services/privacy/pii_scrubber.py — PIIScrubber:
- scrub(text: str) → str
- Replace patterns using regex:
- Borrower/person names: → [PERSON]
- Property addresses (street addresses): → [ADDRESS]
- SSNs (###-##-####): → [REDACTED]
- Loan numbers (10+ digit strings): → [REDACTED]
- Dollar amounts ($xxx,xxx): → [VALUE] (preserves relative language like "above" "below")
- Email addresses: → [EMAIL]
- Return scrubbed text
- Unit tests: verify each pattern is caught, verify dollar amounts lose specific values but preserve context
QUALITY SCORER:
backend/app/services/rules/quality_scorer.py — QualityScorer (0-100 weighted):
1. ComparableQualityScorer (weight 30%):
- Proximity: penalize comps >1 mile urban, >5 miles rural
- Recency: penalize comps >6 months, fail >12 months without explanation
- Physical similarity: GLA within 20%, similar age bracket, similar condition
- Data source: MLS preferred over courthouse records
2. AdjustmentConsistencyScorer (weight 25%):
- Direction: superior comp adjusted down, inferior comp adjusted up — flag inversions
- Bracketing: subject GLA and features bracketed by comp range
- Magnitude: no single line adjustment >10% of sale price
3. MarketAnalysisScorer (weight 20%):
- Market section completeness (all UAD fields populated)
- Trend direction matches neighborhood narrative
- Price range supported by comparable sale prices
4. NarrativeQualityScorer (weight 15%) — Vertex AI powered (see below)
5. ReconciliationScorer (weight 10%):
- Final value within range of comp indicators
- Reconciliation comment explains approach weighting
- Value not rounded to nearest round number matching a single comp
VERTEX AI INTEGRATION:
backend/app/services/rules/quality_narrative_scorer.py:
- Use google-cloud-aiplatform Python SDK
- Model: gemini-2.5-pro (us-central1)
- Init: vertexai.init(project="tf-ai-qc-prod", location="us-central1")
- GenerativeModel("gemini-2.5-pro")
Flow:
1. Run PIIScrubber.scrub() on ALL narrative text before any Vertex call — no exceptions
2. Cache check: SHA256(scrubbed_text) → check in-memory cache (TTL 24 hours)
3. If cache miss, call Vertex AI with this prompt:
System: "You are an expert appraisal reviewer evaluating UAD 3.6 appraisal narrative quality for USPAP compliance and GSE acceptability. Score the commentary on 0-100 based on: specificity (not boilerplate), support for value conclusion, market condition analysis depth, internal consistency, and absence of unsupported statements. Return ONLY valid JSON, no markdown: {\"score\": integer, \"flags\": [{\"field\": \"string\", \"issue\": \"string\"}]}"
User: [scrubbed narrative text]
4. Parse JSON response, validate schema
5. Store in cache with SHA256 key
6. Graceful fallback: if Vertex AI call fails (timeout, quota, error), return score=50 with flag: "Narrative scoring unavailable — manual review required"
7. Log to Cloud Logging: report_id, score, duration_ms, cache_hit (bool), model="gemini-2.5-pro" — NEVER log the text content
CRITICAL: verify that no PII reaches Vertex AI by running tests that confirm scrubber fires before every API call.
Mock Vertex AI in all tests.
SESSION 3D — Rule Admin + Full Pipeline Wiring
You are building TF AI-QC on Google Cloud. Use extended thinking.
Read: docs/superpowers/specs/2026-06-25-uad-qc-tool-design.md
Wire the complete QC pipeline end to end.
1. Seed rules DB — backend/app/db/seed_rules.py:
Insert all implemented rules with rule_code, category, severity, gse_applicability, active=True, config={thresholds}
2. QC Service — backend/app/services/qc_service.py — QCService.run_qc(report_id):
a. Load Report from DB → update status=qc_running
b. Generate 15-min signed URL → download file bytes from GCS (use signed URL, not permanent URL)
c. Delete file bytes from memory after parsing (do not hold in memory longer than needed)
d. Run ExtractorFactory.extract() → ReportData
e. Run RuleEngine.run(report_data) → RuleEngineResult
f. Run QualityScorer.score(report_data) → quality_score (includes Vertex AI narrative scoring)
g. Save QCResult and QCFlag records
h. Update Report status=qc_complete
i. Log to Cloud Logging: report_id, run_number, hard_pass, quality_score, flag_count, vertex_ai_used (bool), duration_ms
j. NOTHING else in the log — no addresses, no appraiser names, no file contents
3. Cloud Tasks — backend/app/services/task_queue.py:
enqueue_qc_job(report_id: str) → creates Cloud Task pointing to POST /internal/run-qc/{report_id}
Queue: tf-ai-qc-prod / us-central1 / qc-jobs
Task deadline: 5 minutes (QC should complete in <2 min normally)
4. Internal routes — backend/app/api/routes/internal.py:
POST /internal/run-qc/{report_id} — secured with Cloud Tasks OIDC token verification
Only callable by Cloud Tasks service account, not by frontend
5. Rule admin API (admin only) — backend/app/api/routes/rules.py:
GET /rules — all rules with active status and config
PATCH /rules/{rule_code} — toggle active, update config thresholds
POST /rules/cache/refresh — force RuleEngine cache reload
6. Results API:
GET /reports/{id}/results — full QC result, flags organized by category then severity
Run full pipeline test: upload sample XML → trigger QC → poll results.
Show QC result JSON output including quality_score and flags.
SESSION 4A — Workflow State Machine
You are building TF AI-QC on Google Cloud. Use extended thinking.
Read: docs/superpowers/specs/2026-06-25-uad-qc-tool-design.md
Build workflow state machine and notification system.
backend/app/services/workflow/state_machine.py — ReportStateMachine:
Valid transitions only:
submitted → qc_running (system only, on Cloud Task start)
qc_running → qc_complete (system only, on QC finish)
qc_complete → approved (reviewer/admin only, requires hard_pass=True AND no open required revisions)
qc_complete → revision_requested (reviewer/admin only)
revision_requested → resubmitted (uploader only)
resubmitted → qc_running (system only, increment run_number, trigger new Cloud Task)
Each transition:
1. Validate caller role can make this transition
2. Validate current status allows transition
3. Update Report.status in DB
4. Write audit log to Cloud Logging: report_id, from_status, to_status, user_id, timestamp
5. Send email notification (if applicable)
6. Raise TransitionError for invalid transitions
Notifications — backend/app/services/notifications.py:
Use SendGrid HTTP API (requests library, no SDK):
- send_revision_requested(appraiser_email, report_id, revision_count): "A revision has been requested on report {report_id}. {revision_count} item(s) need your attention."
- send_resubmitted(reviewer_email, report_id, run_number): "Report {report_id} has been resubmitted (run {run_number})."
- send_approved(appraiser_email, report_id): "Report {report_id} has been approved."
SendGrid API key from Secret Manager (key: sendgrid-api-key)
CRITICAL: email subjects and bodies must NOT contain subject property address, borrower name, or any NPI. Use report IDs only.
API routes (add to reports.py):
POST /reports/{id}/approve — reviewer/admin only
POST /reports/{id}/request-revision — reviewer/admin only (body: initial_message, creates first Revision)
POST /reports/{id}/resubmit — uploader only (multipart: new file — full replacement, not patch)
Integration tests for all transitions including invalid role attempts.
SESSION 4B — Revision Request System
You are building TF AI-QC on Google Cloud. Use extended thinking.
Read: docs/superpowers/specs/2026-06-25-uad-qc-tool-design.md
Build complete revision request and response system.
backend/app/api/routes/revisions.py:
Reviewer endpoints:
- POST /reports/{id}/revisions — create revision
Body: flag_id (optional FK to QCFlag), message, priority: "required"|"recommended"
- GET /reports/{id}/revisions — all revision threads with response history
- PATCH /revisions/{id} — update message or close (reviewer only, must provide close_reason)
- POST /qc-flags/{id}/waive — waive flag (body: waiver_reason required)
Appraiser endpoints:
- GET /my/revisions — all open revisions across all own reports, grouped by report_id
- POST /revisions/{id}/respond — submit response (body: message)
- GET /revisions/{id} — single thread with all responses
Business logic:
- Report cannot be approved while any priority="required" revision is status != "closed"
- Closing a revision requires reviewer or admin role
- Waiving a QC flag requires reviewer or admin + documented reason
- Every revision create/close/waive: log to Cloud Logging (report_id, revision_id, action, user_id) — no message content in logs
Add to GET /reports/{id} response:
- revision_count, open_required_count, open_recommended_count
Integration test: full thread lifecycle — create revision → appraiser responds → reviewer closes → report approved.
SESSION 5A — Appraiser Portal
You are building TF AI-QC on Google Cloud. Use extended thinking for UX decisions.
Read: docs/superpowers/specs/2026-06-25-uad-qc-tool-design.md
Build Appraiser Portal in React/TypeScript/Tailwind.
Auth: Firebase — token auto-attached to all API calls via the axios client from Session 1B.
frontend/src/pages/AppraiserPage.tsx — tabs: My Reports | Upload | My Performance
frontend/src/components/appraiser/ReportQueue.tsx:
- Table: subject address (truncated), form type, submitted date, status badge, quality score (if run)
- Status badges: Submitted=gray, QC Running=blue pulsing, Revision Needed=orange, Approved=green
- Auto-refresh every 30 seconds using @tanstack/react-query refetchInterval
- Click row → ReportDetail drawer slides in from right
frontend/src/components/appraiser/ReportUpload.tsx:
- Drag-drop zone: accepts .xml and .pdf, max 50MB
- Client-side file type detection before upload (don't rely on extension)
- Upload progress bar (Axios onUploadProgress)
- Error states: wrong type, too large, upload failed, server error
- After successful upload: show "QC in progress" status, auto-poll for completion
frontend/src/components/appraiser/ReportDetail.tsx:
- Header: report ID, form type, status badge
- QC Summary section: PASS/FAIL banner (hard compliance), quality score gauge (circular SVG, 0-100, color-coded)
- Flags section: if hard_pass=false, list compliance failures by field. If quality flags, list warnings.
- Revisions section: if status=revision_requested, show RevisionThread for each open revision
- Resubmit button: visible only when status=revision_requested
frontend/src/components/appraiser/RevisionThread.tsx:
- Linked QC flag reference (field name + rule message) if flag_id is set
- Reviewer message with timestamp and reviewer name
- Appraiser response form: textarea + Submit button
- Thread history: all prior messages in chronological order, role-labeled
Use Tailwind only, no component library. Clean functional UI.
SESSION 5B — Reviewer Dashboard
You are building TF AI-QC on Google Cloud. Use extended thinking.
Read: docs/superpowers/specs/2026-06-25-uad-qc-tool-design.md
Build Reviewer Dashboard.
frontend/src/pages/ReviewerPage.tsx — two-panel layout (report list left, review panel right)
frontend/src/components/reviewer/ReportDashboard.tsx:
- Filter bar: status, appraiser name (text search), form type, GSE (FNMA/FHA/VA/FHLMC), date range
- Table: appraiser name, form type, GSE indicators, submitted date, status badge, quality score, open revision count
- Sort by any column
- Click row → loads ReportReviewPanel in right panel
frontend/src/components/reviewer/ReportReviewPanel.tsx:
- Subject info bar: report ID, form type, run number, submitted date
- Compliance section: PASS/FAIL banner. If fail: FlagCard list sorted by severity (fail first)
- Quality section: overall score gauge + 5 category subscores as horizontal progress bars
- Revisions section: grouped by status (Open, Appraiser Responded, Closed)
- Action bar: [Approve Report] | [Request Revision (freeform)]
Approve button disabled when: hard_pass=false OR open_required_count > 0
Disabled state shows tooltip explaining why
frontend/src/components/reviewer/FlagCard.tsx:
- Severity icon: red circle (fail), orange triangle (warning), blue info (info)
- Field name, rule message, GSE reference badge(s)
- Expandable inline form: textarea + priority toggle (Required/Recommended) + Submit as Revision
- Waive button → modal: "Waive flag for [rule_code]?" + required reason textarea + Confirm
- Already-waived flags shown with waiver reason and who waived
frontend/src/components/reviewer/RevisionManager.tsx:
- All revision threads for selected report
- Status chips: Open (orange), Appraiser Responded (blue), Closed (green)
- Click to expand: full message thread + reply textarea (reviewer) + Close button
SESSION 5C — Admin Panel
You are building TF AI-QC on Google Cloud. Use extended thinking.
Read: docs/superpowers/specs/2026-06-25-uad-qc-tool-design.md
Build Admin Panel.
frontend/src/pages/AdminPage.tsx — tabs: Users | Rules | Reporting | Coaching
frontend/src/components/admin/UserManagement.tsx:
- Table: name, email, Google account, role badge, license number, state, active status
- Add User: requires email (must have signed in at least once — Firebase account must exist), name, role, license, state
- Toggle active/inactive with confirmation
- Change role dropdown (appraiser → reviewer → admin, one level at a time)
- Note shown: "User must log in once before role assignment"
frontend/src/components/admin/RuleConfiguration.tsx:
- Table: rule code, category badge, severity badge, description, GSE applicability tags, active toggle
- Toggle active: confirmation dialog "Disabling [rule_code] means reports with [issue] will not be flagged. Continue?"
- Expand row: config JSON editor for threshold values (e.g., net_adjustment_threshold: 15)
- [Refresh Rule Cache] button at top
frontend/src/components/admin/SystemReporting.tsx:
- KPI cards: reports this week, reports this month, first-pass approval rate %, avg quality score, avg turnaround hours
- Bar chart: top 10 most common flags by rule_code (last 30 days)
- Appraiser performance table: name, reports submitted, revision rate %, avg quality score, first-pass rate %
Sortable by each column
- Date range filter (default: last 30 days)
- Use recharts for charts
SESSION 5D — Coaching Dashboards
You are building TF AI-QC on Google Cloud. Use extended thinking.
Read: docs/superpowers/specs/2026-06-25-uad-qc-tool-design.md
Build coaching and performance layer — key product differentiator.
BACKEND first — backend/app/api/routes/coaching.py:
GET /coaching/appraiser/{id}/summary (reviewer/admin only):
Returns: revision_rate (last 30 days), quality_score_trend (array of {month, avg_score}), first_pass_rate, top_flag_categories (top 3 rule categories), vs_team_avg (quality_score_delta)
GET /coaching/team/summary (admin only):
Returns: team_avg_quality, team_avg_revision_rate, first_pass_rate, top_flags (top 10 rule_codes with count)
GET /coaching/patterns (reviewer/admin only):
Returns: appraisers flagged 3+ times in same category in last 30 days
GET /coaching/my/summary (appraiser — own data only):
Returns: own quality_score_trend, category_scores vs anonymous team avg, focus areas (below avg), strengths (above avg)
NEVER return individual appraiser comparisons to appraisers — only vs anonymous aggregate
All coaching queries: parameterized only (no raw SQL string formatting).
FRONTEND:
frontend/src/components/admin/CoachingDashboard.tsx:
- Team KPI cards: avg quality score, avg revision rate, first-pass rate (30-day rolling)
- Recurring Issues panel: "4 appraisers flagged for Adjustment Consistency this month" style alerts
- Appraiser performance table with 6-month quality score sparkline per row
- Click appraiser → AppraiserPerformanceCard modal
frontend/src/components/admin/AppraiserPerformanceCard.tsx:
- Quality score line chart (6 months) — recharts LineChart
- Category radar chart: Comparables, Adjustments, Market, Narrative, Reconciliation vs team avg
- Recent flags list: last 10 flags, rule_code + date + severity
- Private coaching notes field (reviewer/admin only — stored in DB, never shown to appraiser)
frontend/src/components/appraiser/MyPerformance.tsx:
- Own quality score trend chart (6 months)
- Category scores vs anonymized team avg (horizontal bars, "Team avg" as benchmark line)
- "Your strengths" section: top 2 categories above team avg
- "Focus areas" section: bottom 2 categories below team avg
- No ranking, no names, no individual comparison data visible
SESSION 6 — Deploy to Cloud Run + Firebase Hosting
You are deploying TF AI-QC to Google Cloud. Use extended thinking for deployment configuration.
GCP Project: tf-ai-qc-prod | Region: us-central1
Cloud SQL: tfaiqc-db / database: tfaiqc
GCS Bucket: tf-ai-qc-reports-prod
Artifact Registry: us-central1-docker.pkg.dev/tf-ai-qc-prod/tfaiqc-images/backend
Firebase project: tf-ai-qc-prod
BACKEND — Cloud Run:
1. Build Docker image: docker build -t us-central1-docker.pkg.dev/tf-ai-qc-prod/tfaiqc-images/backend:latest .
2. Push: docker push us-central1-docker.pkg.dev/tf-ai-qc-prod/tfaiqc-images/backend:latest
3. Deploy to Cloud Run:
gcloud run deploy tfaiqc-backend \
--image=us-central1-docker.pkg.dev/tf-ai-qc-prod/tfaiqc-images/backend:latest \
--region=us-central1 \
--platform=managed \
--allow-unauthenticated \
--add-cloudsql-instances=tf-ai-qc-prod:us-central1:tfaiqc-db \
--set-secrets=DB_PASSWORD=db-password:latest,SENDGRID_API_KEY=sendgrid-api-key:latest,FIREBASE_ADMIN_KEY=firebase-admin-key:latest,SECRET_KEY=secret-key:latest \
--min-instances=0 \
--max-instances=10 \
--memory=512Mi \
--cpu=1
4. Note the Cloud Run service URL
5. Run Alembic migrations: gcloud run jobs create migrate-db --image=... --command=alembic,upgrade,head
6. Seed rules: gcloud run jobs create seed-rules --image=... --command=python,-m,app.db.seed_rules
FRONTEND — Firebase Hosting:
1. Set API URL: VITE_API_URL=[Cloud Run service URL] npm run build
2. Update firebase.json: add rewrites for /api/* → Cloud Run URL
3. firebase deploy --only hosting
4. Note the Firebase Hosting URL
CLOUD BUILD pipeline (cloudbuild.yaml):
Steps: test backend → build Docker image → push to Artifact Registry → deploy to Cloud Run → run migrations → build frontend → deploy to Firebase Hosting
Trigger: git push to main on github.com/kzelenakas/TF-AI-QC-GCP
POST-DEPLOY SMOKE TESTS:
- ☐ GET {cloud_run_url}/health → 200
- ☐ Firebase Hosting URL loads React app
- ☐ Google sign-in works (uses @truefootage.com account)
- ☐ POST /reports/upload with test XML → returns report_id
- ☐ Cloud Task fires automatically after upload
- ☐ GET /reports/{id}/results returns QC result
- ☐ GET signed URL returns file (expires in 15 min)
- ☐ POST /reports/{id}/request-revision works
- ☐ Email notification sent via SendGrid
- ☐ POST /reports/{id}/approve works after revisions closed
Show Cloud Run service URL and Firebase Hosting URL.
SESSION 7A — Pattern Detection Engine
You are building TF AI-QC on Google Cloud. Use extended thinking.
Read: docs/superpowers/specs/2026-06-25-uad-qc-tool-design.md
Beta is deployed. Build the coaching pattern detection engine.
backend/app/services/coaching/pattern_detector.py — PatternDetector:
- detect_recurring_issues(appraiser_id, lookback_days=30) → list[PatternAlert]
PatternAlert: appraiser_id, rule_code, occurrence_count, first_seen, last_seen, severity_summary
Trigger: same rule_code flagged 3+ times in lookback period
- get_appraiser_trends(appraiser_id) → AppraiserTrends
Monthly avg quality_score, revision_rate, first_pass_rate, top 3 flag categories (last 6 months)
- get_team_benchmarks() → TeamBenchmarks
Anonymous aggregates only: avg quality, std_dev, percentile bands (25th, 50th, 75th)
Results cached 1 hour
- compare_to_team(appraiser_id) → ComparisonResult
Percentile rank for quality_score and revision_rate (no names, no individual rankings)
Alembic migration for coaching_alerts table:
- id, appraiser_id (FK), rule_code, occurrence_count, period_start, period_end, acknowledged (bool), created_at
Nightly Cloud Scheduler job (runs at 2am ET):
- POST /internal/coaching/run-patterns (Cloud Tasks, service account secured)
- For each appraiser: run detect_recurring_issues(lookback_days=30)
- Insert new coaching_alert records for patterns not yet acknowledged
- Monday 8am ET: send weekly digest to all reviewers via SendGrid
All coaching queries: parameterized SQL only. Never log individual appraiser names in Cloud Logging — log only aggregate counts and alert IDs.
SESSION 7B — Coaching Report Export + AI Recommendations
You are building TF AI-QC on Google Cloud. Use extended thinking.
Read: docs/superpowers/specs/2026-06-25-uad-qc-tool-design.md
Build exportable coaching reports and AI-powered training recommendations.
backend/app/services/coaching/report_generator.py — PDFReportGenerator:
- generate_appraiser_report(appraiser_id, period_start, period_end) → bytes (PDF)
Content: name, license, summary period, quality score trend chart (plotted with matplotlib→embed in PDF), category breakdown table, top 3 focus areas with specific rule references, comparison to team avg (anonymized)
- generate_team_report(period_start, period_end) → bytes (PDF)
Content: team KPIs, top 10 recurring issues table, quality score distribution chart
Use reportlab for PDF generation.
backend/app/services/coaching/recommendations.py — RecommendationEngine:
- get_recommendations(flag_categories: list[str]) → list[Recommendation]
Recommendation: issue (category name), resource (guide name), section (specific section), description (what to review)
- Uses Vertex AI Gemini 2.5 Pro:
- PII scrub input: only flag category names sent — no appraiser details, no report content
- Prompt: "You are a USPAP and GSE compliance trainer. For an appraiser who is repeatedly receiving QC flags in these categories: {categories}, recommend specific sections of USPAP, Fannie Mae Selling Guide, or FHA/VA handbooks to review. Return JSON: [{\"issue\": str, \"resource\": str, \"section\": str, \"description\": str}]"
- Cache by sorted(categories) hash (TTL 24 hours)
- Fallback: hardcoded recommendations if Vertex AI unavailable
API:
- GET /coaching/reports/appraiser/{id}?start=&end= → PDF download (reviewer/admin only)
- GET /coaching/reports/team?start=&end= → PDF download (admin only)
- GET /coaching/appraiser/{id}/recommendations → AI-generated training recommendations
Frontend: [Download Coaching Report] button on AppraiserPerformanceCard (admin only).
[View Training Recommendations] panel on both AppraiserPerformanceCard and MyPerformance.
SESSION 8 — Bubble.io OMS Integration
STOP: Do not run Session 8 until you have Bubble.io API documentation and the
bubble-api-tokensecret populated in Secret Manager. Thedata-sources/bubble-oms/folder is currently empty — add Bubble API config files there first.
You are building TF AI-QC on Google Cloud. Use extended thinking.
Read:
- docs/superpowers/specs/2026-06-25-uad-qc-tool-design.md
- data-sources/bubble-oms/ (all files — Bubble API docs and config)
Connect TF AI-QC to the True Footage Bubble.io OMS.
backend/app/services/integrations/bubble_client.py — BubbleClient:
- Auth: API token from Secret Manager (key: bubble-api-token)
- get_order(order_id) → dict with order details
- get_active_orders(appraiser_email) → list[dict] orders assigned to this appraiser
- update_order_qc_status(order_id, status, report_id):
Status mapping: submitted→"QC Submitted", qc_complete+hard_pass=True→"QC Complete", revision_requested→"QC Revision Needed", approved→"QC Approved"
Integration points:
- Report model: add bubble_order_id (nullable) — new Alembic migration
- Upload endpoint: accept optional bubble_order_id in request body
- State machine: on every status transition, call BubbleClient.update_order_qc_status if bubble_order_id is set
- GET /integrations/bubble/my-orders (appraiser only): list active Bubble orders to link on upload
Resilience requirements:
- If bubble-api-token secret is empty, skip ALL Bubble calls silently — system must work without Bubble
- If Bubble API call fails: log the error (order_id, status, error message), do NOT fail the QC workflow
- Failed updates: retry via Cloud Tasks with exponential backoff (3 attempts: 30s, 2min, 10min)
- After 3 failures: create a coaching_alert-style record for manual reconciliation
Frontend: order picker on ReportUpload — dropdown of open Bubble orders for the appraiser to link report to (optional).
SECTION 5 — Cost Estimate (Beta)
| Service | Beta Usage | Est. Monthly |
|---|---|---|
| Cloud Run | 1 vCPU, ~100 req/day | $5–15 |
| Cloud SQL (db-f1-micro) | Always on | ~$10 |
| Cloud Storage | <10GB | ~$0.23 |
| Vertex AI (Gemini 2.5 Pro) | ~200 reports/mo | ~$30–60 |
| Firebase Hosting | Static files | Free |
| Firebase Auth | <10,000 users | Free |
| Cloud Tasks | ~200 QC jobs/mo | <$1 |
| Secret Manager | <10 secrets | <$1 |
| Cloud Logging | Standard retention | Free tier |
| SendGrid | <3,000 emails/mo | Free |
| TOTAL | ~$50–90/month |
Set billing alert at $150/month in GCP Console → Billing → Budgets & Alerts.
SECTION 6 — GLBA Compliance Reference
| Requirement | How TF AI-QC Meets It |
|---|---|
| Encryption at rest | GCS + Cloud SQL encrypt by default (AES-256) |
| Encryption in transit | Cloud Run enforces HTTPS/TLS 1.2+ automatically |
| Access controls | Firebase Auth + role-based (appraiser/reviewer/admin) |
| Signed URLs (15 min) | Enforced in GCSStorageService — not configurable to longer |
| AI data handling | PII scrubber before all Vertex AI calls + Vertex AI model is not trained on customer data (Google DPA) |
| Audit logging | Cloud Logging: all report file access, status changes, sign-ins logged |
| No PII in logs | Enforced in QCService and notifications — only report IDs logged |
| Incident response | 72-hour notification to federal regulator required if NPI exposed — assign owner before launch |
| DPA | Google Cloud DPA accepted when you accept GCP ToS — explicitly review at GCP → IAM → Privacy |
| Data retention | 7-year retention for reports, 3-year minimum for audit logs — implement Cloud Storage lifecycle rules |
SECTION 7 — Pre-Production Security Checklist
- ☐ Firebase Auth restricted to
@truefootage.comdomain - ☐ Cloud Storage bucket has public access blocked
- ☐ All GCS file access via signed URLs with 15-min expiry
- ☐ PII scrubber tested: confirm no names/addresses appear in Cloud Logging
- ☐ PII scrubber tested: send a report with PII-heavy narrative through QC, verify logs show no PII
- ☐ Secret Manager holds all credentials — no .env files on Cloud Run
- ☐ Cloud SQL SSL connections required
- ☐ Audit logging enabled: Data Access logs in IAM → Audit Logs
- ☐ CORS restricted to Firebase Hosting domain only (not wildcard *)
- ☐ Rate limiting active: 100 requests/minute per user
- ☐ Billing alert set at $150/month
- ☐ Google Cloud DPA explicitly accepted
- ☐ Service account has minimum required roles only
- ☐ Integration test: appraiser cannot access another appraiser's reports
- ☐ Integration test: viewer cannot access reviewer endpoints
- ☐ Incident response contact assigned (who gets paged if a breach occurs)
- ☐ Data retention lifecycle rules set on GCS bucket (7 years)
SECTION 8 — Useful GCP Console Links
TF AI-QC Build Guide v2.0 — 2026-06-28
Stack: Google Cloud Run + Cloud SQL + Cloud Storage + Vertex AI (Gemini 2.5 Pro) + Firebase
True Footage internal use only. Review with legal/compliance before production launch.