SWE Task Evaluator - Fully Remote | Upto $90/hr
mercor
Job Description
<h3>About the job</h3><p><strong>Mercor</strong> connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include <strong>Benchmark</strong>, <strong>General Catalyst</strong>, <strong>Peter Thiel</strong>, <strong>Adam D'Angelo</strong>, <strong>Larry Summers</strong>, and <strong>Jack Dorsey</strong>.</p><p><strong>Position:</strong> SWE-Bench Task Auditor<br><strong>Type:</strong><strong>Contract</strong><br><strong>Compensation:</strong><strong>$70–$90/hour</strong><br><strong>Location:</strong><strong>Remote</strong></p><h3>Role Responsibilities</h3><ul>
<li>Evaluate the quality, correctness, and reproducibility of <strong>software-engineering benchmark tasks</strong>.</li>
<li>Assess repository-level tasks, reference patches, test harnesses, and grading integrity.</li>
<li>Provide clear, rubric-based written feedback to improve <strong>AI model training</strong>.</li>
<li>Audit reference patches, test runners, and <strong>Docker</strong> isolation to detect answer leakage and reward hacking.</li>
<li>Work <strong>independently and asynchronously</strong> to meet deadlines and enhance <strong>AI model performance</strong>.</li>
</ul><h3>Qualifications</h3><p></p><p><strong>Must-Have</strong></p><ul>
<li><strong><strong>3+ years</strong> professional <strong>software engineering</strong> experience.</strong></li>
<li><strong>Real open-source contribution or maintainer experience (merged PRs, committer/maintainer roles).</strong></li>
<li><strong>Strong ability to audit reference patches, test runners, and <strong>Docker</strong> isolation.</strong></li>
<li><strong>Fluency across common ecosystems (<strong>Python</strong> and at least one of <strong>Java</strong> / <strong>Go</strong> / <strong>TypeScript</strong> / <strong>C++</strong>).</strong></li>
</ul><h3><strong>Preferred</strong></h3><ul>
<li><strong>Familiarity with <strong>SWE-Bench (Verified)</strong> or similar repository benchmarks.</strong></li>
<li><strong>Maintainer history on major <strong>Python OSS</strong> (<strong>Django</strong>, <strong>Flask</strong>, <strong>scikit-learn</strong>, <strong>sympy</strong>, <strong>pytest</strong>, etc.).</strong></li>
<li><strong>Prior code-review or task-grading experience.</strong></li>
</ul><h3><strong>Application Process (Takes 20–30 mins to complete)</strong></h3><ul>
<li><strong>Upload resume</strong></li>
<li><strong>AI interview based on your resume</strong></li>
<li><strong>Submit form</strong></li>
</ul><h3><strong>Resources & Support</strong></h3><ul>
<li><strong>For details about the interview process and platform information, please check: </strong></li>
<li><strong>For any help or support, reach out to: </strong></li>
</ul><p><strong><em>PS: Our team reviews applications daily. Please complete your AI interview and application steps to be considered for this opportunity.</em></strong></p><p>Originally posted on <a href="https://himalayas.app">Himalayas</a