See llms.txt for all machine-readable content.

Back to Templates

Track new Clutch agencies in Google Sheets using Apify

Created by

Created by: FalconScrape || falcon-scrape
FalconScrape

Last update

Last update 3 days ago

Categories

Share


Quick overview

This manual workflow uses Apify’s Clutch.co Listings Scraper to crawl a single public Clutch directory and writes newly observed agency profiles into Google Sheets, maintaining a “current” view, a per-run history log, and a run ledger with recovery support.

How it works

  1. Runs manually and reads the target Clutch directory URL, maximum items, charge limit, Google Spreadsheet ID, and optional resume run ID.
  2. Validates the configuration and checks the destination Google Sheet, creating the required tabs and a temporary lock to prevent overlapping writes.
  3. Reads and validates existing rows in the Current, History, and Runs tabs to ensure managed headers and safe row limits.
  4. Starts a new Apify Actor run for the Clutch.co Listings Scraper (or loads a specified Apify run to resume) and records the run ID in the Google Sheets run log as a checkpoint.
  5. Polls Apify until the run succeeds, then downloads the complete bounded dataset of scraped agency listings.
  6. Normalizes and deduplicates agencies by Clutch profile slug, compares them to the Current tab, and prepares updates plus appended history observations and run log metrics.
  7. Writes the updated Current view, appends History and Runs entries in Google Sheets, releases the lock tab, and outputs a run summary including the spreadsheet link.

Setup

  1. Add Apify API credentials and Google Sheets OAuth2 credentials for the HTTP requests.
  2. Create an empty Google Sheets spreadsheet and paste its spreadsheet ID into the workflow configuration.
  3. Set the target to a public Clutch directory URL (no query parameters) and adjust maxItems and maxChargeUsd to control dataset size and Apify spend.
  4. (Optional) If a previous execution was interrupted, copy the saved Apify run ID into resumeRunId and ensure the temporary lock tab is removed only after confirming no execution is still running.