Appearance
Status note (2026-09-25): Dated implementation plan. Revalidate current producer/consumer code, payload versions and ticket state before applying its steps.
TikTok/Instagram Submission ScrapeJob Compatibility Bridge
TL;DR: plan-only proposal for bridging scrape jobs across backend and scraper. Never implemented as written — revalidate current producer/consumer code, payload versions, and ticket state before using any of its steps.
Date: 2026-09-06. Status: implementation plan only.
1. Objective and First-Wave Boundary
Merge and deploy the smallest compatibility bridge that removes Apify from TikTok and Instagram initial submission fetching and bio verification. Preserve the current synchronous frontend/API interaction while routing acquisition through the existing database-backed ScrapeJob queue and metric-scraper.
The entire immediate merge path is two coordinated urgent PRs:
- Backend: additive migration, enqueue-and-wait bridge, existing business validation, and route cutover.
- metric-scraper: VIDEO_DETAILS and PROFILE_METADATA handlers for TikTok and Instagram.
There is no separate capability-spike PR, YouTube dependency, SubmissionIntake, frontend work, shared contract package, queue redesign, or application-wide Apify removal in this wave. Stable author IDs and fresh profile bios are implemented and verified inside the scraper PR and deployment smoke tests.
The delivered claim is precise: TikTok/Instagram initial submission creation and their bio-verification checks no longer call Apify; they enqueue ScrapeJobs that metric-scraper claims through the existing queue. Payout refresh, PVTracker, legacy tooling, and other Apify consumers remain.
| Repository | Working location / planned branch | Audited baseline |
|---|---|---|
| Backend | Existing backend worktree; feat/submission-scrape-bridge from dev-wt-1; PR to dev | 226381bdef25f92a77d01f406e993c4c50f9350b |
| Scraper | Existing metric-scraper main worktree; feat/submission-scrape-bridge from main; PR to main | 7fedb346f3fb9520ad380e5b9eaab5791ee97d96 |
The scraper main worktree has pre-existing untracked src/observability/. Do not move, delete, edit, stage, or commit it. Branch creation and bridge edits stay in that main worktree as originally requested.
Non-negotiable correctness:
- Never use a handle as a stable TikTok/Instagram account ID.
- PROFILE_METADATA fetches the current profile bio and returns the observed stable account ID.
- Backend keeps account ownership, bio-code comparison, cooldown/lockout, campaign eligibility, duplicates, moderation, and persistence decisions.
- Worker writes technical job results only.
- No backend fallback to Apify when a platform bridge flag is enabled.
- Validate completed job data strictly in both repositories. Temporary duplicate validators are acceptable and marked for cleanup.
- Timed-out HTTP requests leave durable jobs available for retry reuse.
- Submission and initial ViewSnapshot commit atomically.
Without SubmissionIntake, use the selected/reused VIDEO_DETAILS job createdAt as the durable request time for the 30-minute upload rule. The ten-minute result-reuse window carries that timestamp through the normal verification resubmission. A later request outside the reuse window is a new request. A later asynchronous intake phase will replace this bridge approximation with explicit requestedAt.
2. Current Initial-Submission Call Path
CampaignDetailScreenorSubmitVideoScreenopensfeatures/submissions/components/SubmitVideoModal.tsx.submitClip()validates campaign selection and URL feedback, posts{videoUrl, campaignId}with session cookies, and shows a synchronous checking state.- Backend
src/api/index.ts:314applies write/user rate limits and authentication.requireAuthchecks current user/session revocation and active bans.src/api/routes/submissions.ts:275parsesSubmitVideoSchema: HTTPS and explicit platform host allowlist, positive campaign ID. - The route checks campaign existence/deletion, launch, active state, deadline, pause, accepting-submissions flag, allowed platform, and Instagram Reel URL shape. It checks exact URL duplication within campaign and the same user's ten-minute cross-campaign cooldown.
- It synchronously fetches technical data. YouTube calls
getVideoDetails([id])insrc/utils/youtube.ts. TikTok dynamically importsscrapeTikTok; Instagram dynamically importsscrapeInstagram, both insrc/utils/scrapers/apify.ts. TikTok/Instagram results are cached in process for ten minutes per user/campaign/platform/raw URL. - The route requires a platform video ID. For
requireNewUploads, it requires a valid non-future publish time, strictly after campaign launch, and no more than 30 minutes before the check. It checks canonical duplication by(campaignId, platform, platformVideoId). - It requires a stable creator account ID and looks up
LinkedSocialAccount. Unlinked accounts receive HTTP 412 withrequiresVerification, platform, stable account ID, and handle. Accounts owned by someone else receive 403. No submission exists yet. AccountVerificationModalcalls/api/verifications/start, which creates/reuses a ten-minute bio-code challenge./api/verifications/checksynchronously callsfetchProfileBio: another Apify actor plus dataset read for TikTok/Instagram, or YouTubechannels.list. Backend compares stable ID and bio code, enforces cooldown/lockout, and transactionally creates the link, deletes the challenge, and logs success. Frontend then repeats the submission POST. The in-memory video cache may avoid another video scrape; YouTube is fetched again.- Backend forces TikTok/Instagram to short form, applies the existing YouTube URL/duration heuristic, checks
acceptsShorts, and rejects long form. It creates aPENDINGmoderation submission with initial/current metrics and metadata,nextPollAt=now, then separately attempts the initialViewSnapshot. Snapshot failure is logged but does not undo submission creation. - HTTP 201 contains
{message, submission}. Frontend does not consume the scraped metrics or submission object to render success; any successful HTTP status currently triggers success. A bare 202 would therefore be falsely presented as completed intake unless the frontend changes. - The API's 30-minute tracking tick queues due submissions; a separate 15-second reconciliation loop applies worker results. Tracking now includes
PENDING,ACCEPTED,DENIED, andFLAGGED; it is not gated on moderator acceptance. Cadence/expiry useSubmission.createdAt, notacceptedAt.
There is no initial-fetch pending state today. Submission.status=PENDING means awaiting moderation after technical and ownership checks. PendingVerification represents a bio challenge, not a submission or queued scrape. Submission history fetches on filter/page changes and has no polling mechanism.
Exact Apify Dependency
Both video adapters call getClient() before invoking their actors. It throws when APIFY_TOKEN is absent. The intake route catches that exception and returns generic HTTP 500. The profile adapter independently requires APIFY_TOKEN and returns a transient verification error when unavailable. APIFY_API_KEY is accepted as an alternative only by PVTracker in this codebase. Tests must unset both variable names.
For each unlinked-account submission there can be a video actor call/dataset read, one or more profile actor calls/dataset reads, and another video actor call if cache reuse fails or expires. TikTok uses clockworks/tiktok-scraper; Instagram uses apify/instagram-scraper, both configurable. Actor wait time is 240 seconds. An available local worker cannot help because this route never enqueues its initial fetch.
3. Field-Level Data Audit
TT = TikTok; IG = Instagram; YT = YouTube. Current worker result means data persisted to ScrapeJob.result, not intermediate parser locals. Source files: backend src/utils/scrapers/apify.ts, src/utils/youtube.ts, src/utils/profileBio.ts; scraper src/results/normalize.ts, src/platforms/*/parser.ts, src/platforms/youtube/scrape.ts.
| Field / category | Current initial source | Used for / persisted where | Existing metric worker |
|---|---|---|---|
| Original URL / target | Request videoUrl | Duplicate precheck; Submission.videoLink | url preserves original input |
| Platform / target | Backend URL detection | Campaign platform check; submission and jobs | platform: all three |
| Platform video ID / metadata | TT canonical URL ID mapped to actor row; IG shortcode; YT API item ID | Required; authoritative campaign/platform/video uniqueness | video_id: TT ID, IG shortcode, YT ID; nullable type; IG main response currently lacks identity cross-check |
| Canonical URL / target | Not stored separately | Implicit URL/ID matching | TT/IG helpers compute canonical URLs but result drops them; YT preserves original only |
| Views / metric | TT playCount; IG videoPlayCount, videoViewCount, playsCount; YT statistics.viewCount | initialViews, currentViews, initial snapshot | views: TT embed playCount; IG post or author/coauthor clips fallback; YT API. Nullable contract; old adapters and current YT worker coerce missing values to zero |
| Likes / metric | TT diggCount; IG likesCount; YT statistics.likeCount | initialLikes, currentLikes, snapshot likes | likes: TT embed with player API exact-count supplement, IG post, YT API |
| Comments / metric | TT commentCount; IG commentsCount; YT statistics.commentCount | currentComments, snapshot comments | comments: same respective paths |
| Shares / metric | Not consumed at initial submission | Not seeded today; recurring can populate currentShares and snapshot shares | TT nullable shares; IG/YT null |
| Saves / metric | Not fetched by initial adapter | No current submission/snapshot column | Worker saves; TT universal hydration may supply; TT embed and IG/YT null |
| Title/caption / video metadata | TT text/title/desc/description; IG caption/title/alt/description; YT snippet.title | Optional videoTitle; actor adapter normalizes whitespace and truncates to 500 | Missing on all platforms |
| Preview media URL / video metadata | TT/IG actor videoUrl; YT not assigned | Optional previewVideoUrl | Missing on all platforms |
| Preview image URL / video metadata | TT coverUrl; IG displayUrl; YT not assigned | Optional previewImageUrl | Missing on all platforms |
| Duration / video metadata | TT/IG hard-coded 0; YT ISO 8601 contentDetails.duration | YT short-form heuristic; integer duration | Missing on all platforms; YT requests only statistics,snippet |
| Content kind / technical metadata | IG Reel URL check; TT treated as short | Backend eligibility input | IG parser requires media_type=2, which alone does not prove Reel product type; TT URL helper also accepts photo although intake rejects it; no result discriminant |
isShort / business-derived value | TT/IG true; YT explicit /shorts/ or duration <=180 seconds; unknown duration defaults short | Campaign acceptsShorts and long-form rejection; stored videoType | Missing, and should stay backend-owned as an eligibility classification |
| Creator stable account ID / author metadata | TT authorMeta.id; IG ownerId; YT snippet.channelId | Required ownership/link lookup; not stored on Submission today | Missing from result on all platforms. IG parser has authorId=user.pk, but sink drops it. TT does not parse author ID; YT ignores channel ID |
| Creator handle/display name / author metadata | TT authorMeta.name/uniqueId; IG ownerUsername; YT snippet.channelTitle | Verification prompt and creatorHandle | author_handle available/null; YT value is a display name, not an actual @handle |
| Follower count / profile metric | Not used by initial route | Not required; no write | author_follower_count, TT when exposed; IG/YT null |
| Publish time / video metadata | TT createTimeISO/createTime; IG timestamp; YT snippet.publishedAt | Conditionally required new-upload check; postedAt | posted_at: TT hydration createTime, IG taken_at, YT publishedAt, nullable |
| Stable profile ID + canonical handle / profile metadata | Fresh profile actor: TT authorMeta.id/name, IG id/username; YT channel item id and customUrl/title | Link identity, rename protection, LinkedSocialAccount | No profile operation |
| Bio / profile metadata | TT authorMeta.signature; IG biography; YT snippet.description | Backend matches verification code; not Submission data | Missing; a fresh profile fetch is necessary after user edits bio |
| Hidden-like indicator / technical metadata | YT helper derives likesHidden | Not read by initial route or stored | Not returned; nullable metric should express absence going forward |
| Permission, campaign eligibility, ownership, duplicates, moderation / business results | Auth, campaign rows, linked accounts, backend rules | Determines whether a real Submission is created | Correctly absent; never add to scraper result |
| Observation time and technical diagnostics | Initial snapshot uses backend time; legacy actor result has no observation envelope | Snapshot timestamp; technical queue state | scraped_at, status/error, latency, outbound attempts, Crawlee retries, proxy ID, HTTP status |
Essential differences are stable author identity, title/previews, YT duration, explicit resolved target identity, and fresh profile metadata. Metrics-only cannot satisfy intake.
Evidence Limits and Platform Scope
Current scraper fixtures are mostly synthetic metric responses. Historical sanitized provider fixtures explicitly remove captions, media URLs, profile information, and many identifiers. Their TT author object contains only uniqueId; they cannot prove stable TT account ID or preview extraction. IG historical fixture retains code, pk, user.pk, and username, supporting a proper media identity check.
Use new sanitized fixtures preserving the missing fields with synthetic values. Candidate TT sources are full video hydration itemStruct.author.id, desc, video, and profile hydration; candidate IG sources are the existing media item user.pk, caption, video_versions, image_versions2, plus profile responses. These mappings require parser evidence before being called supported. No live platform availability was tested in this audit.
Keep currently accepted URL classes: TT /@user/video/id and resolve already-allowed vm/vt links in the worker; IG /reel/ and /reels/; YT current watch/share/shorts forms. The worker has requires*Resolution helpers but never calls them, so resolution is actual new work. Do not widen IG intake to posts, /share, or instagr.am just because metric helpers recognize them. Do not accept TT photos as videos. Revalidate every redirect's HTTPS scheme/host and final platform; bound redirects and request timeouts.
4. Current Queue and Result Contract
prisma/schema.prisma:477 defines ScrapeJob: id, required submissionId, platform, url, kind default TRACKING, priority default 0, dueAt, status default QUEUED, attempts, leasedAt, leaseExpiresAt, workerId, scrapedAt, completedAt, JSON result, errorCode, errorMessage, appliedAt, createdAt, updatedAt.
There is no payload, schema version, explicit idempotency key, attempt limit, or retryability column. The two partial unique indexes from 20260831120000_add_scrape_jobs and 20260831213000_hold_tracking_jobs_until_applied enforce active and, more strongly, one unapplied TRACKING job per submission.
Backend src/utils/tracking/scrapeJobs.ts inserts/coalesces with the unapplied partial index. Normal priority is 0; manual sync is 100. Higher-priority requests can make existing active work due now; terminal unapplied results are held for reconciliation. Due selection includes all three platforms and both due and null nextPollAt. After enqueue, backend advances nextPollAt using the existing cadence. Manual sync and queue-all helpers still accept only TT/IG, despite recurring YT support; that unrelated API limitation need not change here.
Worker src/worker/job-repository.ts uses a transaction and FOR UPDATE SKIP LOCKED to claim due QUEUED rows or expired RUNNING leases, currently filtered to TRACKING and three platforms. It orders priority descending then dueAt ascending, increments attempts, and sets a five-minute lease by default. Completion compares job ID, RUNNING status, worker ID, and leasedAt. Stale completions cannot normally overwrite a newer database lease.
Worker polls every five seconds, claims one row per call by default, and continuously feeds long-lived TT/IG Crawlee crawlers (default concurrency 10 each, three request retries). Empty polls sleep; successful claims do not. There is no bound on locally pending leased work, no heartbeat, and no max DB attempts. YT executes one API request inline. Those details can undermine urgent scheduling if the process preclaims a large backlog.
The current metric JSON fields are platform, video_id, url, scraped_at, views, likes, comments, shares, saves, author_handle, author_follower_count, posted_at, status, error, latency_ms, attempts, retries, proxy_id, http_status. attempts in JSON counts outbound requests, unlike the database's lease-claim count. Worker types are TypeScript interfaces, not runtime schemas; historical documentation claiming full worker-side Zod validation does not describe this implementation.
Backend validates only a Zod subset, requiring exact platform and original-URL equality and non-null views on success. It does not validate video ID, complete diagnostics, kind/version, or failed-result JSON. Failure codes are split out of prose by the worker; worker retryability overrides are lost on persistence. unknown and session_error classifications differ across repositories.
Reconciliation scans 100 terminal unapplied TRACKING rows, defaults to every 15 seconds, checks current tracking eligibility, serializes campaign budget calculation with an advisory lock, and commits business changes, ViewSnapshot, and appliedAt together. Only ACCEPTED submissions consume campaign budget; other moderation states still track. Success schedules 12-hour polls for the first 48 hours after creation, then 24 hours. Transient failure uses 24-hour jittered backoff, flags after three failures, and continues tracking; permanent failure flags/stops. A technical failed job is terminal today; the backend schedules another recurring job later.
5. Minimal First-Wave Contract
Keep TRACKING, its version-0 payload/result format, claim ordering, partial indexes, reconciliation, cadence, and failure behavior unchanged.
Add only VIDEO_DETAILS and PROFILE_METADATA for TikTok and Instagram, plus these ScrapeJob fields:
- contractVersion Int with default 0.
- payload nullable JSON.
- dedupeKey nullable String.
- nullable submissionId.
Keep url required. PROFILE_METADATA stores a derived profile URL for diagnostics, while its validated locator payload is authoritative. Do not add intake/check models, resultVersion, lease tokens, retry revisions, heartbeats, max attempts, retention tables, or snapshot-source columns.
Existing rows remain version 0 with null payload/dedupeKey. New rows are version 1. Add a SQL CHECK allowing the legacy TRACKING shape and requiring version 1, null submissionId, non-null payload, and non-null dedupeKey for the new kinds. Add one partial unique index on (kind, dedupeKey) where kind is VIDEO_DETAILS or PROFILE_METADATA, dedupeKey is non-null, and appliedAt is null. Do not alter either TRACKING index.
Implement matching strict Zod schemas in each repository. The duplication is intentionally temporary. Commit identical golden JSON fixtures and a compatibility test; do not add a shared package or deployment dependency.
ts
type BridgeRequestV1 =
| {
contractVersion: 1;
kind: "VIDEO_DETAILS";
platform: "tiktok" | "instagram";
target: { originalUrl: string };
}
| {
contractVersion: 1;
kind: "PROFILE_METADATA";
platform: "tiktok" | "instagram";
target: { expectedAccountId: string; handle: string };
};
type VideoDetails = {
videoId: string;
canonicalUrl: string;
authorAccountId: string;
authorHandle: string | null;
postedAt: string | null;
metrics: {
views: number | null;
likes: number | null;
comments: number | null;
shares: number | null;
saves: number | null;
};
title: string | null;
previewVideoUrl: string | null;
previewImageUrl: string | null;
};
type ProfileMetadata = {
accountId: string;
handle: string | null;
canonicalUrl: string | null;
bio: string;
};
type RedactedProfileMetadata = {
accountId: string;
handle: string | null;
canonicalUrl: string | null;
bioRedacted: true;
};
type BridgeFailure = {
code: BridgeFailureCode;
disposition: "RETRYABLE" | "PERMANENT" | "OPERATOR_ACTION";
message: string;
};
type BridgeResultBase = {
contractVersion: 1;
platform: "tiktok" | "instagram";
observedAt: string;
};
type BridgeResultV1 =
| (BridgeResultBase & {
kind: "VIDEO_DETAILS";
outcome: "SUCCESS";
data: VideoDetails;
})
| (BridgeResultBase & {
kind: "PROFILE_METADATA";
outcome: "SUCCESS";
data: ProfileMetadata;
})
| (BridgeResultBase & {
kind: "PROFILE_METADATA";
outcome: "REDACTED";
data: RedactedProfileMetadata;
})
| (BridgeResultBase & {
kind: "VIDEO_DETAILS" | "PROFILE_METADATA";
outcome: "FAILURE";
failure: BridgeFailure;
});Runtime schemas use separate literal branches so profile and details data cannot cross-validate. Reject unknown fields, invalid dates/URLs, unsafe counts, mismatched kind/platform/target, and missing stable IDs. A fetched empty bio is valid; a missing bio is failure. Validate payload after claim and result before worker write, then validate again after every backend read.
Payloads contain no campaign rules, verification codes, linked-account decisions, or business eligibility. Dedupe keys are SHA-256 hashes:
- VIDEO_DETAILS: version, kind, webUserId, campaignId, platform, and trimmed URL.
- PROFILE_METADATA: version, kind, PendingVerification ID, challenge createdAt, platform, expected account ID, and normalized handle.
Challenge createdAt distinguishes regenerated codes without placing the code in the job. Never store raw provider responses, cookies, credentials, or proxy secrets.
6. Backend Urgent PR
Planned PR: feat/submission-scrape-bridge to dev.
Migration and Queue Bridge
Add only the schema changes above. Build a narrow module with enqueue-or-reuse, bounded wait, strict decode, and consume functions. Enqueue at priority 200 with dueAt equal to database now. The partial unique index is the concurrency guard. Resolve unique conflicts by selecting the winner. Do not hold a transaction/connection while waiting.
Reuse policy:
| Existing work | Behavior |
|---|---|
| Matching QUEUED/RUNNING row | Wait on the same job |
| Matching terminal row with appliedAt null | Consume the same result |
| Successful VIDEO_DETAILS completed within 10 minutes | Reuse for verification resubmission |
| Permanent VIDEO_DETAILS failure within 10 minutes | Reuse to avoid repeated invalid/private extraction |
| Consumed retryable/operator VIDEO_DETAILS failure under 2 seconds old | Reuse and honor Retry-After; later retry may create a new job |
| Consumed PROFILE_METADATA result | Do not treat as a fresh bio check; existing cooldown/lockout precedes a new job |
| Stale unapplied bridge row | Atomically mark closed, then insert a new immutable job |
Automatic Crawlee transport retries remain within one job. Explicit retries create new rows after the old row is consumed. Never reset terminal job data.
Bounded Synchronous Wait
Default SUBMISSION_SCRAPE_WAIT_MS is 25 seconds, clamped to 5-45 seconds. Poll the job primary key from 250 ms up to 750 ms with small jitter. Stop on terminal state, request cancellation, or deadline. A slow/down worker leaves the job queued/running.
Timeout must be non-2xx because the current submission frontend treats all 2xx responses as success. Return HTTP 503, Retry-After: 2, and:
json
{
"code": "SCRAPE_PROCESSING",
"error": "We are still checking this clip. Please try again in a moment.",
"retryable": true,
"retryAfterSeconds": 2
}Verification returns the same 503 with ok:false, transient:true, and message, matching the existing modal branch. Do not return 202 in this wave. Malformed/operator failures also return visible safe 503 responses and log the job ID/classification. Retryable worker failure does not count as a verification failure. Permanent failures map to current submission/verification behavior.
Submission Creation
For TikTok/Instagram only, replace the Apify call with VIDEO_DETAILS enqueue/reuse/wait. YouTube remains unchanged.
Run current cheap validation before enqueue. After success, run the existing verified video ID, campaign upload-age, canonical duplicate, stable linked-account ownership, short-form, and campaign checks. An unlinked account returns the same 412 response and leaves the details result reusable. The post-verification submission request reuses that result.
Create Submission, initial ViewSnapshot, and mark the job applied in one transaction. Use the details metrics as the first observation and preserve existing recurring tracking scheduling without an immediate second bridge fetch. Canonical Submission uniqueness remains the final concurrent duplicate guard. Video ID, stable author ID, and views are required; publish time is conditionally required. Title/previews remain optional. Never manufacture required values.
Bio Verification
For TikTok/Instagram only, replace fetchProfileBio with PROFILE_METADATA enqueue/reuse/wait. YouTube remains unchanged.
Preserve current lockout, cooldown, challenge ownership, and expiry checks. Build the locator from the stored expected ID and handle; never accept replacements from the client and never send the code to the worker. Require returned accountId to equal the pending stable ID before backend code comparison.
Apply a terminal result transactionally with an appliedAt-is-null CAS so concurrent callers create at most one VerificationAttempt/link. On success create LinkedSocialAccount and delete PendingVerification as today. On wrong code record one failed attempt. A concurrent loser returns established linked/cooldown state without another scrape.
Redact the bio in that same transaction, replacing it with a validated redacted PROFILE_METADATA result that retains technical identity/time and bioRedacted:true. Store neither the bio nor code afterward. Timeout does not consume, redact, log an attempt, or delete the challenge; retry reuses the pending job. After a consumed code-not-found result and existing 30-second cooldown, a new check creates a fresh job and fetches a fresh bio.
Flags and Guard
Add SUBMISSION_SCRAPE_JOB_TIKTOK_ENABLED and SUBMISSION_SCRAPE_JOB_INSTAGRAM_ENABLED, initially false for ordered deployment. Each controls both VIDEO_DETAILS and PROFILE_METADATA. When true, queue errors are returned visibly and no Apify fallback is allowed.
Add route tests with legacy fetchers that throw if invoked on enabled branches, plus a focused import/call boundary check. Retain legacy Apify utilities/dependency because payout/PVTracker still use them. Remove disabled route fallback branches in a small post-cutover cleanup PR.
7. Scraper Urgent PR
Planned PR: feat/submission-scrape-bridge to main in the existing scraper main worktree; exclude src/observability/.
Add Zod directly and explicit (kind, platform) dispatch:
| Kind | TikTok | YouTube | |
|---|---|---|---|
| TRACKING | Existing unchanged | Existing unchanged | Existing unchanged |
| VIDEO_DETAILS | New | New | Explicitly unsupported |
| PROFILE_METADATA | New | New | Explicitly unsupported |
Unknown combinations finish with validated permanent unsupported-kind/platform failures. Remove the dispatch assumption that every non-TikTok/non-Instagram case is YouTube, while preserving valid TRACKING behavior.
Reuse existing crawler/session/proxy machinery. Extend extraction only enough for canonical video identity, metrics, stable primary author ID, handle, publish time, optional title/previews, and current profile bio plus profile stable ID. TikTok ID must come from provider video/profile data, never URL handle. Instagram uses the primary media user ID; coauthors used for metric fallback cannot replace it. Safely resolve currently accepted TikTok short links with bounded redirects/timeouts and final HTTPS/platform validation. Do not expand Instagram beyond current Reel intake.
Keep worker polling, lease columns/recovery, priority-desc/dueAt-asc claim order, retries, and batch configuration. Extend the claim predicate/row decoder for the new kinds and nullable submissionId. Do not add queue fairness, heartbeats, lease tokens, attempt ceilings, or capacity architecture.
Completion keeps workerId+leasedAt CAS. Return explicit RETRYABLE, PERMANENT, or OPERATOR_ACTION failures. Parser/config failures must be visible, never rerouted to TRACKING/YouTube.
Implement stable ID and bio parsing in this PR with sanitized real-response fixtures where obtainable and mocked variants for failures. Document field paths/prerequisites beside tests. This is implementation evidence within the urgent PR, not a separate report or a gate involving YouTube.
8. Two-PR Merge and Immediate Cutover
Review both PRs in parallel and link exact counterpart commit/fixture versions. They merge once these focused blockers pass.
Backend blockers:
- Additive migration replays and preserves TRACKING.
- Concurrent enqueue/reuse and terminal-job consumption work.
- Timeout is 503 SCRAPE_PROCESSING, never success.
- Strict variants reject malformed/mismatched data.
- Submission plus snapshot is atomic and uses initial result.
- Stable profile ID/code comparison remains in backend and applies once.
- Bio is redacted after application.
- Enabled TT/IG route tests prove no Apify calls with both variable names absent.
- Existing YouTube and focused tracking/verification tests pass.
Scraper blockers:
- Existing TRACKING jobs still claim/execute/complete.
- All four new routes dispatch explicitly.
- Fixtures prove stable author/profile IDs and current bio fields.
- URL, private/not-found, rate-limit, transport, parser/config, and unsupported failures are explicit.
- Build, typecheck, and focused tests pass.
Deployment:
- Merge both urgent PRs.
- Deploy backend migration/code with flags false.
- Deploy the scraper and confirm its SHA/new handler support.
- In disposable/staging, enable one platform flag, unset APIFY_TOKEN and APIFY_API_KEY, then create a submission and perform a fresh bio-code verification with a controlled TikTok/Instagram account.
- Verify one urgent job per operation, correct stable identity, current bio marker, backend business decisions, one Submission/snapshot, video result reuse, and no Apify logs/calls. Record queue pickup and response latency.
- Stop the worker and prove timeout returns 503; restart and prove the same queued job completes on retry.
- Enable TikTok production immediately after its smoke passes. Enable Instagram immediately after its smoke passes. Neither waits on the other or YouTube.
- If a platform cannot return stable identity/current bio, leave only that flag false and fix its handler in the urgent wave. Never substitute handles or enable an Apify fallback.
- After both are stable, remove the disabled TikTok/Instagram route fallback branches in a small backend cleanup PR.
With available capacity, target claim on the next default five-second worker poll and response within the 25-second wait. Record actual timing. Queue hardening is not a merge prerequisite.
9. Focused First-Wave Tests
Automated tests use mocked/sanitized responses and prohibit live provider/Apify traffic. Controlled live checks occur during deployment smoke.
| Area | Required |
|---|---|
| Migration | Old TRACKING rows/indexes; new nullable ownership/CHECK/index conflict |
| Submission | No Apify; urgent job; queued/running/terminal reuse; timeout 503; malformed result; required stable identity; existing validation |
| Integrity | Atomic Submission/snapshot/application; concurrent canonical duplicate |
| Verification | Stored locator; no code payload; timeout no attempt; identity mismatch; wrong code; one attempt/link; redaction; fresh post-cooldown job |
| Worker | TRACKING regression; four routes; short-link resolution; primary author/profile identity; unsupported YouTube new kinds |
| Failures | Retryable/permanent/operator; private/not-found; block/rate limit; parser/config visibility |
| Smoke | TT and IG submission+verification; reuse; worker restart; no Apify; latency |
Use real PostgreSQL with independent connections for partial uniqueness and CAS/atomicity. Use an explicitly disposable database, never inherited .env data. Run backend Prisma validation/migration replay/client generation/build/focused tests and scraper build/typecheck/tests.
10. Deferred Work
Nothing here blocks the urgent two-PR merge/cutover:
- SubmissionIntake, asynchronous 202 flow, frontend polling/resume, and durable requestedAt.
- Shared contract source/distribution and removal of duplicate validators.
- Rich explicit retry APIs/history.
- Queue capacity/fairness, heartbeat/lease-token, crash attempt ceilings, and load hardening.
- YouTube VIDEO_DETAILS/PROFILE_METADATA migration.
- Payout rescrape, PVTracker, backfills, and operational tooling.
- Broader preview retention and observability work. First-wave profile bio redaction remains mandatory.
11. Remaining Fetch Paths and Delivery Record
| Path after enablement | State |
|---|---|
| TT/IG initial submission fetch | ScrapeJob VIDEO_DETAILS -> metric-scraper |
| TT/IG bio verification | ScrapeJob PROFILE_METADATA -> metric-scraper |
| YouTube initial/profile | Existing backend official API, deferred |
| Recurring TT/IG/YT | Existing TRACKING unchanged |
| Payout refresh | Existing TT/IG Apify and YT API, deferred |
| PVTracker | Existing Apify discovery/profile, deferred |
| Backfills/smoke/legacy adapters | Remain until separately migrated or dead |
Do not claim application-wide Apify removal.
Recorded planning baseline: backend Prisma validation/typecheck pass; 17 focused tracking tests pass; scraper typecheck and 45 tests pass; frontend typecheck passes. No new migration, handler, live smoke, commit, PR, deployment, or flag change has occurred.
Final wave delivery must list both PRs/SHAs, migration, deployed worker SHA/flags, contract fixture match, supported TT/IG fields, test/build results, smoke job IDs/timing, proof both Apify variables were absent, and remaining Apify paths.
Use Conventional Commits with concise bodies explaining the bridge. This revision changes only the plan.