Scoping an Agentic Pilot That Proves Value in Five Days
A good agentic pilot is one sentence, real data, an evaluator, and a cage — finished in five days for $1,500. Here is how to scope it so you learn something true.
Most “agent POCs” fail before the model is chosen. They fail at scope: five jobs, twelve tools, no criteria, sample data that never matches production, and a success metric of “the room clapped.” A pilot that proves value is narrower and meaner.
Spurlock Studios runs agentic pilots at $1,500 · 5 days. You get a working agent on your real data for one job. You keep it either way. The fee credits toward a build. This spoke is how to scope that week so it tells the truth. Parent map: Agentic Systems Operating Manual.
What an AI agent pilot project is (and is not)
Is: a time-boxed build that puts one agentic job into a runnable path with evaluator, sandbox, and escalate — measured on real inputs.
Is not: a strategy workshop with no artifact; a ChatGPT wrapper with your logo; a promise of AGI on Friday; a twelve-integration platform.
If you need architecture across many initiatives, that is the fractional AI CTO shape. If the path is fully known, you may want automation instead — see When Not to Build an Agent.
How to scope an agent POC
Rule 1 — One sentence job
If you need a paragraph, you have two jobs. Examples that fit a week:
- “Classify new support tickets and draft an internal summary with citations or no_match.”
- “Enrich inbound leads with firmographics and write an internal note; do not email.”
- “Turn a meeting transcript into a task list in our tracker with owner guesses, pending human confirm.”
Examples that do not:
- “Own customer support.”
- “Be our sales team.”
- “Replace the ops department.”
Rule 2 — Real data, thin slice
Ten to fifty real examples beat a thousand synthetic ones. Anonymize if you must, but keep the ugly edge cases. Pilots on toy data prove toy performance.
Rule 3 — Criteria before tools
Write pass/fail acceptance lines on day one. No criteria, no pilot — only a demo. Deep dive: Build the Evaluator Before the Agent.
Rule 4 — Minimum tools
Allowlist the smallest set that can complete the sentence. Prefer drafts and internal fields over customer-visible sends. Sandbox rules: tool-use sandboxes.
Rule 5 — Terminal honesty
Define done, escalate, and abort. Cap revisions. Budget the run. A pilot that cannot stop is not production-shaped.
Rule 6 — Success metrics agreed in writing
Pick two or three: golden-set pass rate target, cost per pass ceiling, escalate rate band, human time saved on the sample. “Feels magical” is not a metric.
Five-day shape (what Spurlock Studios actually does)
Day 1 — Contract. Job sentence, criteria, data access, tool list, out-of-scope list.
Day 2 — Evaluator + fixtures. Golden slice, mechanical checks, first fail cases.
Day 3 — Worker + sandbox. State machine thin path: intake → act → evaluate → revise → done/escalate.
Day 4 — Hardening on real cases. Edge cases, cost caps, logging, human gate if needed.
Day 5 — Receipts. Demo on agreed metrics, you keep the agent, build quote from what we saw — not from a fantasy deck.
Timelines assume access lands on day one. Access delayed is the usual reason “five days” becomes eight.
In scope vs out of scope (steal this table)
| In scope | Out of scope for the pilot |
|---|---|
| One job | Multi-department platform |
| One primary system + 1–2 tools | Every SaaS you own |
| Evaluator harness | Perfect model fine-tunes |
| Internal drafts | Autonomous public sends |
| Thin memory fields | Company-wide “brain” |
| Kill switch + revision cap | Fleet multi-tenant billing |
Out-of-scope items can land on the build quote. They should not land mid-pilot as “quick adds.”
Stakeholder roles
- Sponsor — can declare the job sentence and accept metrics
- System owner — grants API/credentials to a sandbox
- Domain reviewer — labels golden cases and judges edge outputs
- Builder — Spurlock Studios for our pilots
Missing domain reviewer is how you discover on day five that “severity” meant something else.
Red flags that the pilot will lie
- Success defined as executive enthusiasm
- Refusal to allow real data
- Insistence on irreversible actions in week one
- Expanding job sentence daily
- No one available to label failures
Decline or rescope. A false-green pilot is worse than no pilot.
After the pilot
Three honest outcomes:
- Ship path — metrics met; quote a build tier to harden and widen.
- Rescope — agent was wrong shape; automation or human process wins.
- Park — value unclear; you still keep the artifact and learning.
All three beat a zombie POC that never decides.
Pricing and next step
Pilot: $1,500 · 5 business days · you keep it · credit toward build. Packaging and FAQs live on /agentic. Start the conversation at /contact?intent=agentic-pilot.
Writing the job sentence (templates)
Use this template:
“Given [trigger], produce [artifact] for [audience], such that [criteria], using [systems], and never [hard no].”
Examples:
- “Given a new Tier-2 support ticket, produce an internal triage summary for the on-call lead, such that severity is enum-valid and citations-or-no_match hold, using Zendesk+help center, and never email the customer.”
- “Given a new inbound lead, produce an enriched internal note for sales, such that firmographic fields are null-safe and sourced, using CRM+enrichment API, and never merge or delete leads.”
If stakeholders cannot agree on the hard no, you are not scoped.
Data access checklist
- Read credentials to staging or a prod read replica
- Written list of fields allowed to write
- PII handling rules
- Rate limits known
- A backup human path if the agent is down
Day-one blockers are almost always here. Send the checklist before the pilot week starts.
Communication during the five days
Daily async note: what passed, what failed, what is blocked. Mid-week scope freeze — no new tools after day two unless something was impossible. End-of-week readout: metrics table, residual risks, build options with costs grounded in what we saw.
Spurlock Studios runs this cadence so sponsors are never surprised on day five. Book at /contact?intent=agentic-pilot; packaging on /agentic; doctrine in the operating manual.
Pricing psychology without the spin
$1,500 is not “cheap AI.” It is a filter: serious enough to grant data access, bounded enough to decide. If a company cannot find $1,500 and a domain reviewer, they are not ready for agents regardless of model hype.
Sample success scorecard (copy/paste)
| Metric | Target | Actual |
|---|---|---|
| Golden-set pass rate | ≥ 85% | |
| Median revisions to pass | ≤ 2 | |
| Cost per passing run | ≤ $X | |
| Escalate rate | 10–25% early is OK | |
| Irreversible actions auto-sent | 0 |
Fill X from finance comfort, not vendor promises.
Scope change protocol
During the five days, new requests go on a parking lot. If a change is required for the job sentence to make sense, swap it for something of equal size — do not grow. Document swaps in the daily note.
What “you keep it” means operationally
You receive the workflow/agent code or runner export as applicable, credentials documentation, criteria doc, and runbook for escalate. You can run without Spurlock Studios. Support after the week is a separate conversation; the pilot credit toward build is stated on /agentic.
How to scope an agent POC with multiple stakeholders
Run a 45-minute scoping call with a shared doc:
- Each person writes a job sentence silently
- Compare and merge to one
- List hard nos
- List systems
- Draft five criteria
- Pick twenty sample IDs for the golden slice
If step 2 fails, do not book engineering days yet.
AI agent pilot project anti-goals
Write anti-goals explicitly: “Not replacing the team,” “Not sending customer email,” “Not building a company brain.” Anti-goals protect the week when excitement spikes mid-build.
Spurlock Studios’ $1,500 · 5-day structure exists to force this clarity. /contact?intent=agentic-pilot
Risk register for the week
List top risks day one: access delay, criterion disagreement, tool rate limits, sample bias. Assign mitigations. Revisit day three. AI agent pilot project success is as much risk management as prompting.
How to scope an agent POC when legal is nervous
Offer drafts-only, staging credentials, redacted traces, and human gates on writes. Bring legal a diagram of the sandbox. Nervous counsel is often unprotected counsel — show the cage.
End state artifacts checklist
- Job contract markdown
- Evaluator criteria + golden slice
- Tool catalog YAML
- State machine table
- Runbook for escalate
- Cost sheet from the week
- Build options with ranges
If artifacts are missing, the week was a demo, not a pilot. Spurlock Studios’ $1,500 · 5-day offer is designed to leave artifacts you keep — /agentic.
Closing note on honesty of scope
How to scope an agent POC is mostly saying no with a smile. One sentence, real data, criteria, cage, five days. An AI agent pilot project that tries to boil the ocean teaches you nothing you can trust. Spurlock Studios priced the pilot at $1,500 so the week stays honest. Book /contact?intent=agentic-pilot when the sentence is ready.
Freeze scope after day two unless the job sentence itself was wrong — growth mid-week is how pilots lie.
FAQ
What makes a good AI agent pilot project?
One sentence job, real data, explicit evaluator criteria, minimal sandboxed tools, revision and budget caps, and written success metrics — delivered as a runnable system, not slides.
How do you scope an agent POC in a week?
Cut to one job, freeze scope, build evaluator first, implement a thin state machine, measure on a golden slice, and refuse irreversible autonomy until scores earn it. Access and a domain reviewer must be available.
Why does Spurlock Studios price the pilot at $1,500?
It is enough commitment to use real data and real criteria, low enough to decide quickly, and structured so the artifact remains yours with credit toward a full build. See /agentic.
Can we pilot multiple jobs in five days?
Not honestly. Sequence pilots or move to a build/fractional engagement. Parallel jobs in one week recreate the scope failure mode.
What do we need ready before day one?
Job sentence draft, sample of real inputs, API access plan, a domain reviewer, and agreement that public sends/refunds stay out of scope unless explicitly negotiated.
How does pilot scope connect to the operating manual?
The pilot installs the minimum viable stack from the manual: evaluator, sandbox, state machine, cost caps, escalate. Platform concerns come after proof.