Appearance
Status note (2026-09-25): Dated integration snapshot and plan. Use current backend/scraper validators, generated fixtures, provenance checks and tests for the live contract.
BloxClips ↔ metric-scraper integration state
TL;DR: how backend and scraper integrate (validators, fixtures, contract) as of 2026-08-31. Background reading — the live contract is the current validators, generated fixtures, provenance checks, and tests.
Inspection date: 2026-08-31
Scope: replace Apify for BloxClips Submission TikTok/Instagram metrics. This is a state-and-plan document; it intentionally makes no runtime change.
Repositories inspected
The outer workspace contains three independent repositories:
| Repository | Role in this assessment |
|---|---|
Bloxclips-backend/ | Express/Prisma BloxClips backend and the current Submission scheduling, history, payout, and Apify paths. |
metric-scraper/ | Local TypeScript TikTok/Instagram metric scraper, including its Crawlee audit and feature-gated Crawlee shell. |
BloxClips-frontend/ | Present in the workspace but not inspected in depth: this integration is backend/database owned and no frontend work is required for the proposed path. |
PVTracker was inspected only enough to separate it from the Submission work. It is not part of the recommended cutover.
Executive finding
Apify remains on the Submission-critical intake and payout-refresh paths, but no longer on accepted TikTok/Instagram tracking. That recurring path now creates backend-owned ScrapeJob rows consumed by the local worker. The original findings and phased plan below are retained as implementation history; the immediately following update is the current runtime state.
The smallest safe production design is therefore:
- Keep scheduling, manual sync decisions, campaign guards, payout/budget rules, and canonical metric writes in BloxClips.
- Add a small
ScrapeJobtable in the backend database for a durable technical request and its terminal raw result. - Run
metric-scraperas one separate worker process which claims jobs, scrapes URLs, and stores only its normalized result/status/error on that job. - Let backend code reconcile a completed job into
SubmissionandViewSnapshotusing the existing budget/payout-aware success/failure code. The scraper never decidesnextPollAt, campaign state, earnings, or payouts.
This is deliberately PostgreSQL-backed; it needs neither Redis nor a separate queue service. The existing Submission.nextPollAt fields are retained as the scheduler's source of truth until (and unless) the new job table replaces the embedded lease.
Implementation update — 2026-08-31
The accepted-Submission tracking cutover described here is now implemented in the two repositories, with PostgreSQL as the only cross-process boundary:
ScrapeJobmigration20260831120000_add_scrape_jobsprovides the durableTRACKINGjob, due time, priority, worker lease, terminal validated JSON result/error, andappliedAt. Follow-up migration20260831213000_hold_tracking_jobs_until_appliedkeeps one unapplied job per Submission, so a manual sync cannot race past a terminal worker result before the backend records itsViewSnapshot.- The backend scheduler dispatches only due accepted TikTok/Instagram submissions to that table; YouTube remains on its existing direct path. A separate backend-owned reconciliation loop runs every 15 seconds by default (
TRACKING_RECONCILE_INTERVAL_MS) so terminal worker results reach product state promptly without changing the 12h/24h polling cadence. It applies a terminal job in the same transaction as the Submission/ViewSnapshot write andappliedAt, so a second reconciliation cannot add a second snapshot. Campaign budget application is serialized per campaign inside that transaction. POST /api/admin/submissions/:submissionId/synccoalesces or prioritizes the normal database job. It does not call a scraper process directly.metric-scraperhaspnpm worker: it atomically claims due/reclaimable jobs, runs the existingCrawleeScrapeRunner/BasicCrawler, and writes exactly one validatedMetricSnapshotresult or stable typed failure. It never writes a Submission, ViewSnapshot, campaign, payout, or schedule field.- Result values use the metric-scraper snapshot contract.
views,likes, andcommentsremain nullable; backend preserves the last known non-null likes or comments and treats a successful row withoutviewsas typedparse_errorrather than inventing zero.
Required worker configuration: explicit SCRAPER_DATABASE_URL (or, only as a fallback, DATABASE_URL), SCRAPER_WORKER_ID, SCRAPER_WORKER_BATCH_SIZE, SCRAPER_WORKER_LEASE_SECONDS, and SCRAPER_WORKER_POLL_INTERVAL_MS, plus the existing scraper runtime/proxy settings. The backend uses its existing DATABASE_URL.
Backend and worker typechecks pass. The production-shaped path has now also been observed against the configured PostgreSQL database: a worker claimed and completed an Instagram job with non-null views/likes/comments, then backend reconciliation updated the Submission and wrote exactly one ViewSnapshot; no completed jobs remained unapplied afterwards. The full scraper suite retains three pre-existing sandbox failures that require opening a local HTTP listener; focused worker tests pass.
Current Apify usage map
Submission-critical
| Location | Current use | Why it is in cutover scope |
|---|---|---|
Bloxclips-backend/src/utils/scrapers/apify.ts | scrapeTikTok(urls) and scrapeInstagram(urls) invoke Apify actors in batches and normalize views, likes, comments, title, preview URLs, creator data, and posting time. | Shared adapter for the Submission paths below. |
Bloxclips-backend/src/api/routes/submissions.ts | On POST /submissions, TikTok/Instagram are scraped synchronously before the Submission is created. The result initializes current metrics and supplies the account identity used by the social-account ownership/verification flow. | This is an Apify dependency before an accepted submission even exists. |
Bloxclips-backend/src/utils/tracking/runTrackingTick.ts | Dispatches due accepted TikTok/Instagram Submissions to ScrapeJob and reconciles terminal jobs; YouTube remains direct. | The local-worker recurring tracking path. |
Bloxclips-backend/src/utils/payouts/rescrape.ts | A user payout request re-scrapes eligible Submission URLs through Apify before calculating payout items. | Submission/payout-adjacent. It should be deliberately deferred or routed through the same job path after tracking cutover; it must not silently remain a second metric source. |
Bloxclips-backend/src/utils/tiktok.ts | Legacy single-TikTok wrapper around the shared Apify adapter. | Submission/admin helper; check real callers before removal. |
Bloxclips-backend/backfill_initial_metrics.ts, scripts/smokeApify.ts, scripts/backfillApifyTracking.ts | One-off/backfill or Apify smoke tooling. | Not production routing, but cleanup/documentation candidates after cutover. |
Bloxclips-backend/src/utils/scrapers/costTracker.ts and package.json | Apify-cost accounting and apify-client dependency. | Can remain temporarily for non-Submission Apify paths; do not remove globally in the Submission cutover. |
Related but not the Submission video-metric replacement
| Location | Current use | Cutover treatment |
|---|---|---|
Bloxclips-backend/src/utils/profileBio.ts | Apify retrieves TikTok/Instagram profile bios for social-account bio-code verification. | Separate account-profile capability. Do not block Submission metric tracking on replacing it; plan a later local profile adapter if the initial submission/verification flow requires full Apify removal. |
Bloxclips-backend/src/utils/pvTracker.ts | Direct Apify HTTP actor calls discover/profile-sync TikTok and Instagram videos into file-backed PVTracker state. | Explicit non-goal. Leave unchanged. |
Bloxclips-backend/src/utils/tracking/pollScheduler.ts | Comments and terminal-reason matching refer to Apify's current free-form errors. | Not an Apify call site, but its string-based failure contract must be replaced with typed error codes at cutover. |
There is no evidence that Apify currently powers a separate Prisma SubmissionMetric, VideoMetric, or campaign-metrics table. Its Submission results land directly in the Submission row and ViewSnapshot history.
Current Submission metric flow
Creation / review path
POST /api/submissionsvalidates the URL/platform and duplicate rules insrc/api/routes/submissions.ts.- YouTube uses the YouTube API. TikTok/Instagram synchronously call
utils/scrapers/apify.ts(with a brief in-memory cache for the account-verification retry). - The endpoint requires
creatorAccountId, checksLinkedSocialAccount, and can returnACCOUNT_VERIFICATION_REQUIREDbefore creating the submission. - It creates a
PENDINGSubmissionwith current views/likes/comments and writes the firstViewSnapshot. - Admin acceptance in
src/api/routes/admin.tssetsacceptedAtandnextPollAtto now. That is the current equivalent of enqueueing the first scheduled scrape.
Accepted Submission tracking path
src/api/index.tsstartsstartTrackingScheduler()in the API process.- The scheduler invokes
runTrackingTick()every 30 minutes. - The tick atomically leases due rows by updating
Submission.claimedAtandclaimStaleAtusing PostgreSQLFOR UPDATE SKIP LOCKED. Eligibility is accepted,nextPollAt <= now, not stopped/frozen, unclaimed or stale, and on an active unfrozen campaign. - It separates YouTube, TikTok, and Instagram. TikTok/Instagram batches go to Apify.
- On success, backend code applies the campaign budget clamp, writes
currentViews,currentLikes,currentComments, optional metadata,lastPolledAt, the next cadence, and oneViewSnapshotin a transaction. - On failure, backend code increments failures, calculates retry/terminal behavior, updates
nextPollAtor flags/stops the Submission, then releases the lease.
The current recurring flow is durable enough to recover a crashed API tick, but it couples claiming, external scrape execution, business-rule application, and result persistence in one API-process function. It is not consumable by the local scraper as-is.
Database and queue state
Existing models relevant to this integration
Bloxclips-backend/prisma/schema.prisma contains:
| Model / fields | Current role |
|---|---|
Submission | Product source of truth. Stores URL/platform/status, current and initial views, current likes/comments, metadata, campaign relation, payout-related fields, and the tracking fields below. |
Submission.acceptedAt, lastPolledAt, nextPollAt | Backend-owned cadence state. nextPollAt has an index with status. |
Submission.trackingStoppedAt, trackingStoppedReason, consecutiveScrapeFailures, lastMetricsChangedAt | Backend-owned failure/terminal/change state. |
Submission.claimedAt, claimStaleAt | Existing atomic claim lease. The migration 20260831075300_add_submission_claim_and_change_tracking_fields added these fields. |
ViewSnapshot | Append-only-ish per-poll history: submissionId, viewCount, nullable likes/comments, and snapshotDate; indexed by submission/timestamp. |
Campaign | Holds tracking duration and budget/freeze state that the backend applies to scrape results. |
LinkedSocialAccount | Ownership verification record; initial synchronous submission scraping needs its stable account ID. |
There are still no SubmissionMetric or VideoMetric models. ScrapeJob now provides the general durable queue and its normalized terminal JSON result.
Does a durable scrape queue exist?
Yes. ScrapeJob is the durable per-scrape queue for accepted TikTok/Instagram tracking. The legacy Submission due-work fields remain the backend schedule source of truth, and still drive the direct YouTube path:
nextPollAtmakes work due;claimedAt/claimStaleAtimplement a 20-minute visibility lease;- the claim is an atomic
UPDATE … WHERE id IN (SELECT … FOR UPDATE SKIP LOCKED); - successful and failed processing release the lease;
- expired leases are reclaimable without a reaper;
- retry is represented by moving
nextPollAt, and terminal failure changes the Submission state.
The ScrapeJob row adds the request identity/result state that the Submission row lacked: job ID, attempts, worker lease, queued/running/completed status, validated result/error, and idempotent reconciliation marker. Workers do not claim Submission rows directly.
Smallest additional database contract
Add one ScrapeJob model owned by the backend. It can be intentionally narrow:
| Field | Purpose |
|---|---|
id (UUID), submissionId, platform, url | Durable identity and scrape input. submissionId links the request to BloxClips product state. |
kind (TRACKING initially; reserve SUBMISSION_INTAKE only when the pre-create flow is migrated) | Makes the initial implementation explicit without campaign semantics in the scraper. |
priority, dueAt, status (QUEUED, RUNNING, SUCCEEDED, FAILED) | Backend-created due work; manual sync raises priority/sets dueAt=now. |
attempts, leasedAt, leaseExpiresAt, workerId | Claim/reclaim/retry mechanics. |
scrapedAt, completedAt, result JSON or explicit nullable views, likes, comments, optional metadata | Normalized terminal scraper result. Critical counts must be nullable or non-negative integers, never fabricated zeros. |
errorCode, errorMessage | Typed terminal result/error, safe to show/log. |
appliedAt (or a unique result-application record) | Idempotently separates worker completion from backend application to Submission/ViewSnapshot. |
Indexes: (status, dueAt, priority), (leaseExpiresAt), and (submissionId, status) are enough initially. Add a partial unique rule for one active TRACKING job per submission if the chosen migration model requires coalescing; otherwise an idempotency key such as submission:<id>:tracking:<scheduled-period> prevents duplicate polling jobs.
The existing nextPollAt remains backend schedule state. Initially, the scheduler creates or coalesces a TRACKING job when a Submission is due and advances nextPollAt only when the backend applies that job's terminal result. This prevents a worker failure from losing the due scrape and keeps manual sync on the same durable path.
Current metric-scraper capability
Entry points and execution
- CLI:
pnpm cli tiktok <file>,pnpm cli instagram <file>, andpnpm cli run <config>accept batches;--watchrepeats an in-process batch schedule. - Library-level contract:
src/core/scraper/scraper.tsexposesScraper.scrape(url, context): Promise<ScrapeResult>for one attempt. The platform implementations return normalized data or typed failure status/error. - Output:
MetricSnapshotis Zod-validated and containsplatform,video_id, URL, timestamp, nullable views/likes/comments and optional data, terminalstatus/error, and retry/timing fields. The production sink currently isJsonlFileSink;SnapshotSinkis a useful narrow adapter boundary. - Normal execution:
src/cli/execute-batch.ts,src/app/run-service.ts, andsrc/app/scrape-session.tsalways call the legacy in-processScrapeRunner, even when Crawlee flags make a separate runner available. - Crawlee:
crawlee@3.18.1is installed.CrawleeScrapeRunneris a tested, feature-gated/disabled-by-defaultBasicCrawlershell with ephemeral local storage, per-jobuniqueKey, and one terminal snapshot per job. It is not wired into normal CLI/web production routing and does not claim DB work or persist DB results. - Database worker:
pnpm workeruses a narrow PostgreSQL repository, atomicFOR UPDATE SKIP LOCKEDclaim/lease, and aSnapshotSinkthat writes only a claimed job's terminal normalized result. It usesCrawleeScrapeRunnerdirectly; it has no Prisma dependency and never writes BloxClips business tables.
The scraper is now a narrow BloxClips technical worker. Database design and all business scheduling remain outside it; its only durable responsibility is claim → Crawlee → result.
Contract drift
| Backend/Apify expectation | metric-scraper current contract | Required alignment |
|---|---|---|
Batch functions return one result per input URL with a free-form reason. | Single-attempt Scraper.scrape returns typed status/error; runner turns it into MetricSnapshot. | Define a shared worker-result DTO with typed errorCode and nullable metrics; stop terminal logic from matching Apify prose. |
Apify success coerces unavailable metrics to 0. | Scraper intentionally uses null for unknown values. | Preserve null; only views/likes/comments proven available are written. Existing non-null Submission counters can keep their last known values on a partial/failure result rather than turning unknown into zero. |
| Backend calls Apify synchronously during submission creation, before a Submission exists. | Scraper exposes reusable modules but no supported CJS/package/API boundary for the backend, and no synchronous single-URL service entrypoint. | Treat intake/ownership verification as a distinct migration seam. Either create a pre-submission intake job/state or add a small, versioned local scraper client boundary; do not make manual sync direct RPC. |
runTrackingTick owns claim → scrape → campaign clamp → write snapshot. | Scraper owns technical acquisition/retry/output only. | Split backend dispatch/reconciliation from acquisition. Scraper stores raw normalized job result; backend applies business rules. |
nextPollAt and claim lease live on Submission. | CLI --watch and p-queue are process-local, not durable business scheduling. | BloxClips owns due times and manual priority in PostgreSQL. Do not use Crawlee request storage or --watch as the durable scheduler. |
Backend needs current Submission state and history. | JSONL is the current durable output. | Implement a Postgres SnapshotSink/job-result adapter or equivalent worker repository. JSONL can stay for local diagnostics only, not the source of truth. |
One accuracy caveat must be made explicit before the cutover: the scraper README says TikTok may expose source-reported/rounded large view counts even though likes/comments use a second endpoint for exact values. The requested critical field is views, so fixture tests must establish the accepted accuracy behavior for the target URLs before turning on the Submission route. Do not hide this mismatch by rounding or inventing values.
Missing glue
Backend
- Extract the business-rule portions of
runTrackingTickinto reusable result-application functions. They retain campaign-budget clamp, snapshots,nextPollAt, failures, and state changes, but no longer call Apify. - Add
ScrapeJobmigration/repository and a scheduler that creates/coalesces jobs for due accepted Submissions. Add a manual sync endpoint that creates/reprioritizes a due-now job rather than invoking a scraper process directly. - Add a reconciliation loop (can run in the existing API scheduler) that applies each completed job exactly once using the extracted backend functions.
- Replace the Submission route's direct Apify import only after deciding how the pre-creation scrape and account-verification response are represented. This is not solved by an accepted-submission tracking job alone.
- Route payout rescrape deliberately: either enqueue a high-priority
PAYOUT_REFRESHjob through the same queue or leave that call explicitly deferred for the first Submission tracking cutover. Never let it quietly remain the authoritative metric path.
metric-scraper
- Add a production worker entrypoint, separate from CLI/watch/dashboard, which uses the existing normalized scraper contract and is configured with only worker operational settings plus a backend DB connection/repository.
- Add a Postgres job repository or a narrowly scoped backend-owned package/API adapter: claim due jobs with
SKIP LOCKED/leases, persist one validated terminal result, and release/retry leases. It must not import campaign/payout code. - Add a DB
SnapshotSink/result mapper that validatesMetricSnapshotand writes job result fields. Keep JSONL for local CLI usage. - Make Crawlee selection a worker concern only after its feature-gated runner has a production-compatible result callback. The current shell is useful but does not itself satisfy the worker boundary.
Boundary and deployment
The immediate lowest-friction boundary is a separate Node worker deployment from the metric-scraper repository, sharing only the BloxClips PostgreSQL database through a restricted worker DB role (or a small internal backend repository package). A direct TypeScript source import is not appropriate today: the backend is CommonJS/ts-node, metric-scraper is ESM, the scraper is not published as a package, and its composition has different dependencies and lifecycle. Do not add a public HTTP scraping service merely to implement sync-now; the DB queue is the contract.
Required configuration initially:
- backend: existing
DATABASE_URL; a feature flag such asSUBMISSION_LOCAL_SCRAPER_ENABLED; optional dispatch/reconciliation interval and job lease duration; - worker:
DATABASE_URLor a narrowly scoped DB connection string,SCRAPER_WORKER_ID, claim/batch/lease settings, plus the existing scraper proxy, session, rate, and timeout configuration; - rollback: a backend
SUBMISSION_SCRAPE_PROVIDER=apify|localswitch for the brief controlled cutover. KeepAPIFY_*only while unrelated PVTracker/profile work or rollback requires it.
Recommended phased implementation plan
Phase A — isolate Submission Apify usage
Goal: make the boundaries explicit without behavior change.
- Classify
apify.tsconsumers as Submission tracking, Submission intake, payout refresh, profile verification, and PVTracker. - Extract a provider-neutral
SubmissionScrapeResultDTO and typed failure code mapping; retain the Apify adapter behind it temporarily. - Keep PVTracker and profile-bio paths unchanged.
Likely files: Bloxclips-backend/src/utils/scrapers/apify.ts, new src/utils/scrapers/types.ts, src/utils/tracking/runTrackingTick.ts, src/utils/tracking/pollScheduler.ts, src/api/routes/submissions.ts, src/utils/payouts/rescrape.ts.
Verification: backend TypeScript build; focused unit tests for typed terminal/retry mapping; existing scripts/smokeTrackingTick.ts against a test/stub provider, never live Apify in CI.
Phase B — establish DB job and result contract
Goal: add the small durable ScrapeJob table and retain BloxClips ownership.
- Add Prisma schema/migration and job repository/claim SQL.
- Extract
runTrackingTick's success/failure application to accept the normalized job result and atomically writeSubmission+ViewSnapshot+ jobappliedAt. - Dispatch due Submissions into coalesced
TRACKINGjobs. Do not have the scraper changenextPollAt.
Likely files: Bloxclips-backend/prisma/schema.prisma, new migration, src/utils/tracking/runTrackingTick.ts (split), new src/utils/tracking/scrapeJobs.ts, src/utils/tracking/scheduler.ts, and focused tests.
Verification: npx prisma validate; migration on a disposable PostgreSQL database; tests for claim exclusivity, stale lease reclaim, coalescing, exactly-once application, success, retryable failure, terminal failure, and snapshot creation.
Phase C — give metric-scraper a DB worker mode
Goal: one worker claims only technical jobs and records one normalized terminal result.
- Add a DB repository/worker entrypoint and result mapper; use
MetricSnapshotSchemaas its input validation boundary. - Start with the currently stable legacy scraper runner if it is simpler; select Crawlee only through an explicit worker flag after the runner supports the worker result sink.
- Do not enable scraper
--watch, dashboard, campaign logic, payout logic, or direct manual-sync RPC.
Likely files: metric-scraper/package.json, new src/worker/index.ts, src/worker/job-repository.ts, new DB result sink/mapper under src/infrastructure/output/, src/app/composition.ts, and worker tests. The existing src/infrastructure/crawlee/crawlee-scrape-runner.ts is a candidate integration point, not a drop-in DB worker.
Verification: pnpm test -- tests/crawlee tests/platforms; new worker repository tests with PostgreSQL/test container or a transaction-isolated test database; pnpm typecheck; pnpm build.
Phase D — backend dispatch and manual sync
Goal: all recurring and manual accepted-Submission work enters the same DB queue.
- Have the backend create/coalesce due jobs for new accepted Submissions and periodic due polls.
- Add/route a submission “sync now” command to set or create a high-priority due-now job. It returns job state; it does not call the worker directly.
- The backend reconciliation loop applies completed results and sets the next poll date.
Likely files: Bloxclips-backend/src/api/routes/admin.ts (or a new submissions sync route), src/utils/tracking/scheduler.ts, src/utils/tracking/scrapeJobs.ts, and src/api/index.ts only if a separate reconciler startup hook is needed.
Verification: API integration tests show acceptance and sync-now create one coalesced job; worker completion produces current metrics and ViewSnapshot; concurrent sync requests do not produce duplicate active jobs.
Phase E — route Submission metrics away from Apify
Goal: turn on local worker output for accepted TikTok/Instagram Submission tracking.
- Enable the local provider behind a backend feature flag for tracking jobs only.
- Add fixture parity tests requiring views, likes, and comments to match expected values; reject/record missing values rather than defaulting to zero.
- Decide and implement the distinct initial-submission/account-verification seam. The simplest consistent implementation is a short-lived backend-owned intake record/job whose result drives the existing verification response; it is still DB-queued, but need not block Phase D's accepted-tracking rollout.
- Move payout refresh only when its job-kind/application semantics are tested.
Likely files: Phase A–D files plus src/api/routes/submissions.ts, metric-scraper/tests/platforms/*, tests/crawlee/*, and new backend integration tests.
Verification: end-to-end test from due Submission → worker job → reconciled snapshot; critical-field fixture parity; failure/retry/lease recovery; controlled staging run with no duplicate metric writes.
Phase F — remove/disable Submission Apify path
Goal: remove only the Submission use of Apify after the local path proves stable.
- Delete or disable Apify imports in Submission tracking and migrated intake/payout paths.
- Retain Apify dependency/config only if PVTracker, profile bio, or rollback still needs it. Remove cost tracking only when no remaining caller uses it.
Likely files: src/utils/tracking/runTrackingTick.ts, src/api/routes/submissions.ts, src/utils/payouts/rescrape.ts if migrated, src/utils/tiktok.ts, then eventually src/utils/scrapers/apify.ts, costTracker.ts, package.json, and old smoke scripts.
Verification: rg -n 'scrapeTikTok|scrapeInstagram|apify' shows no Submission-critical callers; backend build/tests; worker tests; a staging DB job processed without any APIFY_* credential.
Phase G — leave unrelated Apify uses alone
PVTracker and profile-bio scraping remain independently scoped. Their continued Apify use does not block the Submission metric cutover unless an environment/dependency cleanup would accidentally remove credentials or shared package support they still require.
Explicit non-goals
- No dashboard, monitoring panel, proxy/session telemetry, or rich observability work.
- No PVTracker redesign or migration in this pass.
- No campaign, payout, earnings, or
nextPollAtbusiness logic in metric-scraper. - No Redis, external queue, distributed orchestration system, or Crawlee storage used as product durability.
- No direct worker RPC as the main implementation of “sync now”.
- No big-bang migration of all Apify consumers.
- No manufactured zero for unavailable values, and no claim of exact TikTok views without fixture evidence.
- No Apify compatibility layer beyond a short rollback flag if it is operationally needed.
Commands used for this assessment
Read-only inspection included repository file inventories; full-text searches for Apify, Prisma models, scheduling/claims, worker/DB code, and entrypoints; and source review of the files named above. Recommended implementation verification commands are listed in each phase. No live scraper, Apify, migration, or deployment command was run.