A video capture request should answer one question: what exactly must happen at the location, and what proof will show that it happened? HumanTask API turns that question into a dispatchable task whose purpose is to dispatch a human to record short, task-specific video evidence at a real-world location.
When to Use Video Capture
The best video capture jobs are concrete enough that two independent workers would understand the same objective. Examples include record a walkthrough of a property exterior or permitted interior; capture how a machine, sign or display looks while operating; show traffic flow or queue conditions during a defined window; or record a short event or venue observation for a remote team. These are not broad consulting assignments. They are observable field actions that can be accepted, completed and checked.
Define a Video Capture Request
A requester should provide address or venue, required sequence of scenes, and minimum or maximum duration. It should also state audio requirements, visit time, and orientation and stabilization instructions. The instruction should separate facts the worker can observe from decisions the AI or business will make later. That distinction prevents a simple field task from quietly turning into specialist advice.
Required inputs
- Address or venue
- Required sequence of scenes
- Minimum or maximum duration
- Audio requirements
- Visit time
- Orientation and stabilization instructions
Evidence That Makes the Result Useful
For video capture, a useful completion package may contain original video files, clip labels, capture time, location context, and worker observations about anything the camera could not show clearly. The evidence schema should be selected before dispatch. A task that merely says “send proof” is weaker than one that specifies which files, fields and observations are mandatory. The receiving agent can then test completeness without interpreting a chat message.
Suggested result fields
- Original video files
- Clip labels
- Capture time
- Location context
- Worker observations about anything the camera could not show clearly
Failure Modes to Plan For
Real locations create exceptions that software APIs rarely face. In this capability, common examples are: audio is restricted, the worker cannot enter the requested area, the subject is not present, and the recording window changes at the venue. Those outcomes should be returned as explicit exception states rather than hidden inside a free-text note. An AI agent can then retry with new instructions, choose another location, widen the deadline or escalate to a specialist.
Exception examples
- Audio is restricted
- The worker cannot enter the requested area
- The subject is not present
- The recording window changes at the venue
Example Workflow for an AI Agent
Consider this workflow: A remote facilities agent can ask for a 90-second walkaround of approved equipment, including indicator lights and cable connections, before deciding whether an engineer must travel. The agent first decides that a physical check is necessary, then creates a video_capture task with the address, deadline and evidence fields. The worker accepts the job, completes only the permitted actions and submits the requested proof. After the result arrives, the software can validate required fields, store the media references and continue its original plan.
Who Uses Video Capture
Video Capture can support property teams, event researchers, remote operations, retail auditing and AI systems that need motion or environmental context. These users have different business goals, but they share the same bottleneck: the missing fact or action exists offline. A reusable task definition lets them solve that bottleneck without maintaining a field team in every city.
API Shape for Video Capture
The machine-facing representation should be narrow. A video_capture request can carry a normalized location, human-readable instructions, a deadline, budget, required evidence and a client reference. The response should return a task identifier and lifecycle state. Follow-up operations should expose status and evidence without forcing the caller to scrape a dashboard. MCP can present the same operation as an agent tool; REST and OpenAPI can serve conventional application integrations.
Example task object
{
"capability": "video_capture",
"location": {"address": "TARGET_ADDRESS"},
"deadline": "ISO_8601",
"instructions": "TASK-SPECIFIC_INSTRUCTIONS",
"evidence_required": ["TASK_SPECIFIC_FIELDS"]
}Launch a Video Capture Task
Start with one video capture request that has an unambiguous outcome. Define the location, deadline and proof first; then create the task through the web flow or the available developer interface. If the workflow repeats, promote the same evidence schema into a reusable integration.
What makes video capture different from a generic gig
The value is not simply that a person is available. The value is that the request is standardized enough for software to understand the expected result. For video capture, the schema should reflect the actual decision being supported: the agent needs evidence about dispatch a human to record short, task-specific video evidence at a real-world location. That makes the task easier to price, route, compare and audit than an open-ended message to a freelancer.
{
"task_id": "tsk_example",
"capability": "on_site_photos",
"status": "evidence_submitted",
"evidence": [
{ "type": "photo", "captured_at": "…", "location_confirmed": true }
],
"exceptions": []
}Frequently asked questions
What should a video capture request contain?
Include address or venue, required sequence of scenes, minimum or maximum duration, and audio requirements. Add the remaining task-specific fields when they affect access, proof or timing.
What does HumanTask API return for video capture?
A result can contain original video files, clip labels, capture time, and location context, plus explicit notes when the task cannot be completed as planned.
What can prevent a video capture task from completing?
Typical blockers include audio is restricted, the worker cannot enter the requested area, and the subject is not present. The worker should report the blocker rather than invent a successful result.
Can an AI agent create video capture programmatically?
Yes. The intended machine-facing capability is `video_capture`, using the same task object whether the caller comes through REST, OpenAPI or MCP.
When is video capture a poor fit?
It is a poor fit when the request is unsafe, requires unverified professional expertise, depends on private access that has not been arranged, or cannot be evaluated with observable evidence.
Related pages
- CapabilitiesBrowse real-world tasks that AI agents and businesses can delegate to verified humans worldwide.Explore
- On-Site PhotosHire verified people worldwide to capture on-site photos with task instructions, timestamps and structured evidence.Explore
- Event AttendanceSend a local human to attend an event, capture observations and return structured evidence.Explore
- Equipment InspectionSend a verified human to inspect equipment, capture evidence and answer structured task questions.Explore
- VerificationHumanTask API verifies real-world task completion with structured evidence such as photos, video, timestamps and location data.Explore
- MCPConnect AI agents to real-world human workers through the HumanTask API MCP server.Explore
- LocationsBrowse HumanTask API coverage by country and city for real-world tasks completed by verified humans.Explore
Next step
Put a human on it.
Describe the place, the action and the proof you need. The API and MCP integration are in developer preview.