Skip to content
HumanTaskAPI

Capability · data_collection

Collect Real-World Data with Humans

Deploy verified humans to collect structured data from the physical world for AI and business workflows.

Evidence returnedExample

Data Collection · field capture

captured_at · location_confirmed

Structured checklist

required fields · exceptions

Task result object

capability: "data_collection"

Digital agents are strong at planning and data processing; data collection begins where those abilities stop. The service is designed to deploy humans to gather repeatable, structured observations from the physical world across one or many locations. Instead of opening a generic freelance project, the requester creates a specific task with acceptance criteria and a structured result.

When to Use Data Collection

The best data collection jobs are concrete enough that two independent workers would understand the same objective. Examples include collect labels or prices; record visible attributes at points of interest; capture training images under a protocol; or repeat the same field form across locations. These are not broad consulting assignments. They are observable field actions that can be accepted, completed and checked.

Define a Data Collection Request

A requester should provide data schema, sampling rules, and locations. It should also state field definitions, quality checks, and media requirements. The instruction should separate facts the worker can observe from decisions the AI or business will make later. That distinction prevents a simple field task from quietly turning into specialist advice.

Required inputs

  • Data schema
  • Sampling rules
  • Locations
  • Field definitions
  • Quality checks
  • Media requirements

Evidence That Makes the Result Useful

For data collection, a useful completion package may contain structured rows, photos linked to records, timestamps, location fields, and quality or exception flags. The evidence schema should be selected before dispatch. A task that merely says “send proof” is weaker than one that specifies which files, fields and observations are mandatory. The receiving agent can then test completeness without interpreting a chat message.

Suggested result fields

  • Structured rows
  • Photos linked to records
  • Timestamps
  • Location fields
  • Quality or exception flags

Failure Modes to Plan For

Real locations create exceptions that software APIs rarely face. In this capability, common examples are: sampling instructions are ambiguous, the field definition changes between locations, data would require private or restricted access, and the collection method creates bias. Those outcomes should be returned as explicit exception states rather than hidden inside a free-text note. An AI agent can then retry with new instructions, choose another location, widen the deadline or escalate to a specialist.

Exception examples

  • Sampling instructions are ambiguous
  • The field definition changes between locations
  • Data would require private or restricted access
  • The collection method creates bias

Example Workflow for an AI Agent

Consider this workflow: A vision team can define a collection protocol once and request the same labeled observations in multiple cities as coverage becomes available. The agent first decides that a physical check is necessary, then creates a data_collection task with the address, deadline and evidence fields. The worker accepts the job, completes only the permitted actions and submits the requested proof. After the result arrives, the software can validate required fields, store the media references and continue its original plan.

Who Uses Data Collection

Data Collection can support AI labs, research teams, mapping projects, retail analytics and operations. These users have different business goals, but they share the same bottleneck: the missing fact or action exists offline. A reusable task definition lets them solve that bottleneck without maintaining a field team in every city.

API Shape for Data Collection

The machine-facing representation should be narrow. A data_collection request can carry a normalized location, human-readable instructions, a deadline, budget, required evidence and a client reference. The response should return a task identifier and lifecycle state. Follow-up operations should expose status and evidence without forcing the caller to scrape a dashboard. MCP can present the same operation as an agent tool; REST and OpenAPI can serve conventional application integrations.

Example task object

json · Example
{
  "capability": "data_collection",
  "location": {"address": "TARGET_ADDRESS"},
  "deadline": "ISO_8601",
  "instructions": "TASK-SPECIFIC_INSTRUCTIONS",
  "evidence_required": ["TASK_SPECIFIC_FIELDS"]
}

Launch a Data Collection Task

Start with one data collection request that has an unambiguous outcome. Define the location, deadline and proof first; then create the task through the web flow or the available developer interface. If the workflow repeats, promote the same evidence schema into a reusable integration.

What makes data collection different from a generic gig

The value is not simply that a person is available. The value is that the request is standardized enough for software to understand the expected result. For data collection, the schema should reflect the actual decision being supported: the agent needs evidence about deploy humans to gather repeatable, structured observations from the physical world across one or many locations. That makes the task easier to price, route, compare and audit than an open-ended message to a freelancer.

json · Example result shape
{
  "task_id": "tsk_example",
  "capability": "on_site_photos",
  "status": "evidence_submitted",
  "evidence": [
    { "type": "photo", "captured_at": "…", "location_confirmed": true }
  ],
  "exceptions": []
}

Frequently asked questions

What should a data collection request contain?

Include data schema, sampling rules, locations, and field definitions. Add the remaining task-specific fields when they affect access, proof or timing.

What does HumanTask API return for data collection?

A result can contain structured rows, photos linked to records, timestamps, and location fields, plus explicit notes when the task cannot be completed as planned.

What can prevent a data collection task from completing?

Typical blockers include sampling instructions are ambiguous, the field definition changes between locations, and data would require private or restricted access. The worker should report the blocker rather than invent a successful result.

Can an AI agent create data collection programmatically?

Yes. The intended machine-facing capability is `data_collection`, using the same task object whether the caller comes through REST, OpenAPI or MCP.

When is data collection a poor fit?

It is a poor fit when the request is unsafe, requires unverified professional expertise, depends on private access that has not been arranged, or cannot be evaluated with observable evidence.

Next step

Put a human on it.

Describe the place, the action and the proof you need. The API and MCP integration are in developer preview.