home Home / How to Run a Low-Risk Robot Pilot Before a Full Purchase: Scope, Metrics, and Acceptance Criteria

How to Run a Low-Risk Robot Pilot Before a Full Purchase: Scope, Metrics, and Acceptance Criteria

A pilot is not a demo. A demo shows what a robot can do under controlled conditions. A pilot tests whether a robot can do what your operation actually requires under real conditions, with real parts, real operators, and real constraints. The difference between a well-designed pilot and a poorly designed one is the difference between a confident purchase decision and an expensive mistake. This article explains how to design a robot pilot that reduces purchase risk by testing the right things, measuring the right metrics, and defining what “success” means before the pilot starts.

What a Pilot Should Prove—and What It Cannot Prove

A pilot should prove that the robot can perform the target task at the required level of performance in the buyer’s actual operating environment. Specifically, the pilot should answer:

  • Can the robot handle the actual parts (dimensions, weight, surface, orientation)?
  • Can the robot achieve the required throughput or cycle time?
  • Can the robot operate reliably over a sustained period (not just a single cycle)?
  • Can the robot integrate with the required systems (machine, conveyor, WMS)?
  • Can operators run the robot after training, without constant supplier support?
  • What are the failure modes, and how recoverable are they?

A pilot cannot prove:

  • Long-term reliability (months or years of operation)
  • Multi-robot fleet behavior (if only one robot is piloted)
  • Full-scale throughput (if only a subset of the process is automated)
  • Total cost of ownership (if hidden costs like spares and training are not tracked)
  • Scalability to other sites (if the pilot site is not representative)

Understanding these limits prevents the buyer from over-interpreting pilot results. A successful pilot means the robot is worth purchasing—not that it will definitely succeed at full scale.

Select a Representative Task, Not the Easiest Demo

A recurring pilot design risk is selecting the easiest task to demonstrate, rather than the most representative task for the operation. Suppliers naturally prefer to demo their robot on a task it performs well, but the buyer needs to test the robot on the task it will actually perform.

What Makes a Task Representative

  • The part is typical of the production mix (not the lightest or easiest to handle)
  • The cycle time requirement is at or near the production target (not relaxed for the pilot)
  • The environment is representative (not a clean test area, but the actual production floor)
  • The integration is real (connected to the actual machine, conveyor, or WMS—not simulated)
  • The operator is a typical production operator (not a supplier engineer)

What to Avoid

  • “Lab conditions” pilot: the robot runs in a clean test area with perfect lighting, flat floor, and no traffic. The pilot succeeds, but the production environment is different.
  • “Best case” pilot: the robot handles the easiest part, at the slowest speed, with a supplier engineer operating it. The pilot succeeds, but production requires harder parts at higher speed with a trained operator.
  • “Single cycle” pilot: the robot performs one cycle successfully. The pilot is declared a success, but sustained operation reveals thermal issues, battery limitations, or integration instabilities.

Define Baseline Performance Before the Robot Arrives

Before the robot is deployed, the buyer should measure and document the current performance of the manual or semi-automated process. This baseline is essential for two reasons:

  1. It provides a comparison point: “the robot achieved X, compared to the manual baseline of Y”
  2. It prevents the “pilot succeeds but we don’t know if it’s better” problem

Baseline Metrics to Document

MetricWhat to MeasureWhy It Matters
ThroughputUnits processed per hour (manual)Robot performance compared against current process
QualityDefect rate, rework rate (manual)Robot should not introduce new quality issues
Cycle timeTime per unit (manual)Robot’s cycle time compared against current process
Operator timeHours of operator attention per shiftWhether robot reduces operator time, not just cycle time
Error/recovery rateHow often does the manual process require intervention?Robot’s intervention rate compared against current process
DowntimeHow much time is lost to manual process issues?Whether robot reduces or increases downtime

Note: whether the robot must meet or exceed manual throughput depends on the buyer’s business case. Some projects accept different throughput in exchange for safety, labor availability, consistency, or other operational value. The baseline defines the comparison; the business case defines what “good enough” means.

Choose Metrics: Throughput, Coverage, Task Success, Intervention, Quality and Uptime

The pilot should measure metrics that directly answer “is this robot good enough for production?” The pass criteria for each metric should be buyer-defined, derived from the baseline, the business requirement, and agreed acceptance criteria—not from generic industry thresholds.

MetricDefinitionHow to MeasurePass Criterion
ThroughputUnits processed per hour (or area cleaned per hour, or deliveries per hour)Count completed tasks over a defined measurement period (e.g., 1 shift)Buyer-defined, based on business requirement
Task success ratePercentage of tasks completed without manual interventionCount successful tasks / total tasksBuyer-defined
Intervention rateManual interventions per 100 tasksCount interventions / total tasks x 100Buyer-defined
QualityDefect rate or quality score after robot operationInspect output against quality standardBuyer-defined, compared to baseline
UptimePercentage of scheduled operating time the robot was available(Scheduled time - downtime) / scheduled timeBuyer-defined
Coverage (cleaning)Percentage of target area actually cleanedMeasure cleaned area / target areaBuyer-defined
Recovery timeTime from fault to resumed operationMeasure fault-to-recovery time for each incidentBuyer-defined

Setting Pass Criteria

Pilot acceptance thresholds should be defined before the test based on the buyer’s business requirements, baseline, application criticality, and agreed gap-closure plan. The pilot is a first deployment—the robot, the integration, and the operators are all new. Setting the bar at full production target may reject a robot that would reach full performance after optimization. The specific threshold is a buyer decision based on application criticality, available alternatives, and the gap-closure plan for full deployment.

Test Real Obstacles, Real Parts, Real Operators and Real Peak Conditions

The pilot must test the robot under conditions that reflect actual production, not idealized conditions:

Real Obstacles

  • Production floor with actual layout (racking, machines, pedestrian routes)
  • Dynamic obstacles (forklifts, people, other equipment in motion)
  • Floor conditions as they are (not freshly cleaned or leveled for the pilot)
  • Lighting as it is (not supplemented with extra lights for the pilot)

Real Parts

  • Actual production parts (not samples or mock-ups)
  • Full range of part variations (if parts vary in size, weight, or surface, test the range)
  • Parts in the actual presentation method (conveyor, pallet, bin, fixture)
  • Parts with actual condition (oily, wet, dusty—not cleaned for the pilot)

Real Operators

  • Production operators who will actually run the robot (not engineers or supplier staff)
  • Operators who have received the standard training (not advanced or extended training)
  • Operators running the robot through a full shift (not just a 30-minute demo)

Real Peak Conditions

  • Test during peak production periods (not during low-demand periods)
  • Test with the actual upstream and downstream process running (not in isolation)
  • Test with the actual shift schedule (including shift changes and breaks)
  • Test with the actual network and utility conditions (not a dedicated pilot network)

Document Exceptions and Manual Intervention

Every exception and manual intervention during the pilot should be documented—not to criticize the robot, but to understand its real-world behavior:

Exception Log

Data PointWhy It Matters
What happenedDescribes the failure mode
When it happenedIdentifies patterns (time of day, shift, production phase)
What the robot was doingIdentifies the task or condition that triggered the exception
What the operator didDescribes the recovery action
How long it tookMeasures the recovery time and its impact on throughput
Root cause (if known)Identifies whether the issue is robot, integration, environment, or operator
Was it repeatableDetermines if it’s a systematic issue or a one-time event

This exception log is the most valuable output of the pilot. It tells the buyer what to expect in production and what to plan for in terms of SOPs, training, and support. Review the exception log with the supplier before the purchase decision—patterns in the log reveal whether issues are fixable (integration, SOP) or fundamental (capability mismatch).

Decide Pass / Modify / Stop / Scale

At the end of the pilot, the project team should make one of four decisions:

Pass (Proceed to Full Purchase)

  • All pilot metrics meet or exceed the buyer-defined pass criteria
  • Exception log shows no critical or systematic issues
  • Operators can run the robot without constant supplier support
  • The business case holds at the actual pilot performance level

Modify (Proceed After Changes)

  • Most metrics meet criteria, but specific issues need to be addressed
  • Issues are understood and fixable (e.g., gripper adjustment, route optimization, SOP update)
  • A plan exists to address the issues before full deployment
  • The business case still holds with the expected post-modification performance

Stop (Do Not Proceed)

  • Critical metrics do not meet criteria
  • Issues are not understood or not fixable within budget
  • Operators cannot run the robot after training
  • The business case does not hold at the actual pilot performance level

Scale (Proceed to Multi-Site Rollout)

  • Pilot passes all criteria
  • Configuration is documented and standardized
  • Site readiness process is defined and tested
  • Support model is validated

The “Modify” decision is a frequent outcome—pilots often reveal issues that need to be addressed before full deployment. The key is to distinguish between issues that are fixable (gripper, route, SOP) and issues that are fundamental (robot cannot handle the part, throughput is structurally insufficient).

Pilot Test Checklist for First-Time Buyers

Pilot ElementStatusNotes
Representative task selected (not easiest demo)  
Baseline performance documented (before robot)  
Metrics defined with buyer-defined pass criteria  
Real parts, real environment, real operators  
Peak conditions tested (not just low-demand periods)  
Exception log maintained throughout pilot  
Sustained operation tested (multiple shifts, not single cycle)  
Integration tested (real machine/WMS, not simulated)  
Operators trained and tested (without supplier support)  
Failure modes documented and understood  
Business case validated at actual pilot performance  
Decision made: Pass / Modify / Stop / Scale  
Pilot report documented with metrics, exceptions, and decision  

Request a Quote or Technical Evaluation

Tell us what you need the robot to do. Even if some technical details are not yet confirmed, our team can help evaluate suitable options.

Please share, if available: application, key requirements, site and integration conditions, quantity, destination, and target timeline.

Send Your Requirements

Research Sources Used

  1. Source / organization: EDDIE Project (i-PARIHS framework for pilot-to-scale methodology) | URL: as cited in report_batch_a/b | Version/date: as cited
  2. Source / organization: IEEE (pilot validation and scalability methodology) | URL: https://ieeexplore.ieee.org/ | Version/date: as cited in report_batch_a

Internal product/material source: report_batch_a (batch A research report) [TO VERIFY]: none

Contact Us