Mercor

Mercor San Diego, CA
Job Description Job Description About the job Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark , General Catalyst , Peter Thiel , Adam D'Angelo , Larry Summers , and Jack Dorsey . Position: Biology PhD Coding Experts Type: Contract Compensation: $70/hour Location: Remote Duration: 6 weeks Commitment: 20+ hours/week Role Responsibilities Source material from published papers, Kaggle datasets, or open-source repositories to create research problems. Write scientific prompts and develop grading criteria for correct answers. Calibrate tasks against frontier models, ensuring strong models fail more often than they succeed. Utilize Python , R , or other programming languages for scientific computing. Work with Git/GitHub and Docker for code execution and quality checks. Collaborate asynchronously and independently to meet deadlines and...

Mercor New York, NY
Job Description Job Description About the job Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark , General Catalyst , Peter Thiel , Adam D'Angelo , Larry Summers , and Jack Dorsey . Position: Public Sector Practice Reviewer Type: Contract Compensation: $50–$55/hour Location: Remote Role Responsibilities Evaluate the realism of AI-generated tasks in public sector roles to ensure they reflect actual job scenarios. Assess the accuracy of domain content, verifying that standards, figures, and regulations are consistent and correct. Review AI grading on completed tasks, identifying discrepancies and providing detailed feedback. Determine the defensibility of scores, comparing AI-generated scores with your professional judgment. Work independently and asynchronously to meet deadlines while enhancing AI model performance . Qualifications Must-Have...

Mercor San Francisco, CA
Mercor is recruiting PhD and Master’s level scientists to author AI evaluation tasks (Sci Code) for a new benchmark in scientific computing. You will craft original, executable research problems that current frontier models cannot solve. Requirements include a PhD in biology or related fields, demonstrated depth in two biology subdomains, and proficiency in Python or R. Experience with Docker and Git/GitHub is preferred. Duration is 6 weeks, part-time with 20+ hours per week, and immediate start. #J-18808-Ljbffr

Mercor San Diego, CA
Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code) Mercor is partnering with leading AI labs on a new benchmark for scientific computing. You will author original, executable research problems that today's frontier models cannot solve. Domains — depth required in at least two subdomains (with a coding focus) Biology — ecology, biochemistry, genetics What you'll do Source your own material: a published paper, a Kaggle dataset, an open-source repository, or a scenario you design Write scientific prompts based on the input Build the grading criteria that define a correct answer Calibrate against frontier models — a task ships only when strong models fail it more often than they succeed Required PhD in biology, biological sciences, biochemistry, genetics, ecology, or a closely related field Demonstrated depth in at least two of the following subdomains: ecology, biochemistry, genetics Working proficiency in Python, R, or another relevant...

Mercor New York, NY
Job Description Job Description About the job Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark , General Catalyst , Peter Thiel , Adam D'Angelo , Larry Summers , and Jack Dorsey . Position: Certified Medical Coder (ICD-10) — Multilingual Clinical Documentation & AI Evaluation Type: Contract Compensation: $45–$65/hour Location: Remote Commitment: 10+ hours/week Role Responsibilities Review clinical documentation for accuracy, completeness, and coding integrity. Evaluate AI-generated clinical notes against real-world documentation and coding standards. Annotate clinical documentation against detailed guidelines and provide structured feedback. Validate alignment between clinical documentation and assigned codes. Contribute expert input on annotation guidelines and documentation standards. Work independently and asynchronously with...

Mercor New York, NY
Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code) for a benchmark in scientific computing. You will source material from papers, datasets, or open repositories and design scenarios for evaluation. Responsibilities include writing prompts, building grading criteria, and calibrating tasks so frontier models fail more often than they succeed, with a strong emphasis on biology domains. #J-18808-Ljbffr

Mercor New York, NY
Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code) Mercor is partnering with leading AI labs on a new benchmark for scientific computing. You will author original, executable research problems that today's frontier models cannot solve. Domains — depth required in at least two subdomains (with a coding focus) Biology — ecology, biochemistry, genetics What you'll do Source your own material: a published paper, a Kaggle dataset, an open-source repository, or a scenario you design Write scientific prompts based on the input Build the grading criteria that define a correct answer Calibrate against frontier models — a task ships only when strong models fail it more often than they succeed Required PhD in biology, biological sciences, biochemistry, genetics, ecology, or a closely related field Demonstrated depth in at least two of the following subdomains: ecology, biochemistry, genetics Working proficiency in Python, R, or another relevant...

Mercor New York, NY
Mercor is hiring certified medical coders to support a healthcare AI partner building advanced clinical documentation tools. You will apply your coding and documentation expertise to review, annotate, and evaluate clinical documentation and AI-generated clinical notes, directly shaping the accuracy and reliability of medical AI systems. This role is remote and part-time, requiring at least 10 hours per week. #J-18808-Ljbffr