About the role
Welo is an AI services company focused on training and evaluating AI systems. The Welo Data team builds datasets and quality assessments that help improve AI voice agents. You'd join as a contractor evaluating how well these conversational systems perform in real interactions.
This role involves listening to recorded conversations and rating how an AI voice agent responds. It's part-time work starting immediately, with a commitment of around 10 hours per week. You'll be assessing audio interactions remotely from New Zealand.
Your day-to-day work would look like this
- Listen to recorded conversations between users and AI voice agents, paying close attention to each exchange
- Rate each agent response using two separate 1–5 scales: one for content quality (helpfulness, relevance, appropriateness) and another for prosody (naturalness, pacing, expressiveness)
- Write brief, specific justifications explaining the reasoning behind each score you assign
You'll need strong listening skills and the ability to make consistent judgments across many samples. Written communication matters here—your explanations need to be clear and concise so others can understand your assessment logic. You should be reliable and available for the roughly 10 hours weekly that this role requires.
Nice to have: Experience rating or evaluating audio, familiarity with AI voice systems, or background in linguistics or speech analysis.
What We Offer
- Compensation of $22.10 per hour
- Flexible freelance arrangement with remote work from New Zealand
- Project-based engagement starting as soon as possible
Pay, location & hours
$22 / hour. Fully remote, open to applicants in New Zealand.
About Welo

22 open roles in this building · Company page → · See it on the map