Quick overview
This workflow runs a customer-support AI agent against Google Sheets test cases, uses Groq-hosted LLMs to evaluate response quality, logs results back to Google Sheets, compares run metrics to an approved baseline, and sends regression or review alerts via Gmail.
How it works
- Starts when you manually execute the workflow.
- Reads test cases from Google Sheets, filters to only Active scenarios, and iterates through them in batches.
- Builds a customer-support prompt for each test case and generates an agent response using Google Gemini.
- Sends the question, expected facts/outcome, and the agent response to a Groq LLM to score relevance, reference accuracy, completeness, instruction compliance, and hallucination risk as JSON.
- Parses the evaluator JSON and appends per-test results (including the agent response and overall result) to an Evaluation_Results sheet in Google Sheets.
- Aggregates the current run’s results, selects the latest previously reviewed baseline from Google Sheets, and computes pass-rate regression plus metric deltas.
- Writes the run-level metrics back to Google Sheets and routes the outcome to either mark the run successful or send a Gmail warning/alert to the QA team.
Setup
- Create a Google Sheets file with Test_Cases, Evaluation_Results, and Baseline_Metrics tabs (including columns referenced in the workflow like Active, Test_ID, Question, Expected_Facts, Expected_Outcome, and Reviewed).
- Add Google Sheets credentials in n8n and replace YOUR_GOOGLE_SHEET_ID in all Google Sheets nodes with your spreadsheet ID.
- Add credentials for Google Gemini and Groq (used for the agent and evaluator chat models) and select the models you want to run.
- Add Gmail credentials and update the recipient address ([email protected]) and any email content/subjects to match your QA process.
- Ensure at least one baseline row in Baseline_Metrics is marked Reviewed=true so the workflow can compare the current run against an approved baseline.