Quick overview
This manual workflow uses Apify’s Clutch.co Listings Scraper to crawl a single public Clutch directory and writes newly observed agency profiles into Google Sheets, maintaining a “current” view, a per-run history log, and a run ledger with recovery support.
How it works
- Runs manually and reads the target Clutch directory URL, maximum items, charge limit, Google Spreadsheet ID, and optional resume run ID.
- Validates the configuration and checks the destination Google Sheet, creating the required tabs and a temporary lock to prevent overlapping writes.
- Reads and validates existing rows in the Current, History, and Runs tabs to ensure managed headers and safe row limits.
- Starts a new Apify Actor run for the Clutch.co Listings Scraper (or loads a specified Apify run to resume) and records the run ID in the Google Sheets run log as a checkpoint.
- Polls Apify until the run succeeds, then downloads the complete bounded dataset of scraped agency listings.
- Normalizes and deduplicates agencies by Clutch profile slug, compares them to the Current tab, and prepares updates plus appended history observations and run log metrics.
- Writes the updated Current view, appends History and Runs entries in Google Sheets, releases the lock tab, and outputs a run summary including the spreadsheet link.
Setup
- Add Apify API credentials and Google Sheets OAuth2 credentials for the HTTP requests.
- Create an empty Google Sheets spreadsheet and paste its spreadsheet ID into the workflow configuration.
- Set the target to a public Clutch directory URL (no query parameters) and adjust maxItems and maxChargeUsd to control dataset size and Apify spend.
- (Optional) If a previous execution was interrupted, copy the saved Apify run ID into resumeRunId and ensure the temporary lock tab is removed only after confirming no execution is still running.