

Mercor
Mercor
LLM Red Team Specialist
LLM Red Team Specialist

US

Contract-based

Date Posted

Offered salary
$60 - $90 per hour
$60 - $90 per hour

Closing date
Closing soon
Closing soon


Qualification
MSc / PhD in a stem field
MSc / PhD in a stem field


Hiring location
US
US


Experience
1+ Years
1+ Years
Responsibilities
• Explore how frontier AI models behave on coding, ML, and analysis tasks and find where they quietly get things wrong
• Turn discovered weaknesses into well-crafted tasks that are hard for models but fair to grade
• Write up findings clearly with evidence and reproducible steps
• Team up with task authors to close loopholes, shortcuts, and grading gaps
• Share insights with researchers and fellow experts to improve the benchmark
Requirements
• MSc or PhD in a STEM field, or equivalent practical experience in a research-heavy domain requiring data analysis and coding
• 1+ years of experience in a research, research-engineering, security, or AI-evaluation role
• Demonstrated ability to identify vulnerabilities, edge cases, or failure modes in LLMs or ML systems (via red teaming, adversarial testing, security research, or rigorous model evaluation)
• Working proficiency in Python and Git, with the ability to script your own probes and analyses
• Strong familiarity with LLM capabilities, limitations, and evaluation techniques
• Perfectionist mindset with high attention to detail, creativity in finding what others missed, strong written communication, and ability to work independently on ambiguous problems
• Ability to engage reliably for ~35 hours per week
Preferred
• Past experience in AI training, model evaluation, or benchmark/task authoring
How to Apply
Click "Apply" to be taken to the Mercor website. Complete your profile by uploading your resume and confirming your work location. Once verified, you will be matched to opportunities as they arise. Please note that this role cannot support H1B or STEM OPT candidates. Applying through our link supports WFH Bulletin as a referral partner, but you are welcome to apply directly if you prefer.
Responsibilities
• Explore how frontier AI models behave on coding, ML, and analysis tasks and find where they quietly get things wrong
• Turn discovered weaknesses into well-crafted tasks that are hard for models but fair to grade
• Write up findings clearly with evidence and reproducible steps
• Team up with task authors to close loopholes, shortcuts, and grading gaps
• Share insights with researchers and fellow experts to improve the benchmark
Requirements
• MSc or PhD in a STEM field, or equivalent practical experience in a research-heavy domain requiring data analysis and coding
• 1+ years of experience in a research, research-engineering, security, or AI-evaluation role
• Demonstrated ability to identify vulnerabilities, edge cases, or failure modes in LLMs or ML systems (via red teaming, adversarial testing, security research, or rigorous model evaluation)
• Working proficiency in Python and Git, with the ability to script your own probes and analyses
• Strong familiarity with LLM capabilities, limitations, and evaluation techniques
• Perfectionist mindset with high attention to detail, creativity in finding what others missed, strong written communication, and ability to work independently on ambiguous problems
• Ability to engage reliably for ~35 hours per week
Preferred
• Past experience in AI training, model evaluation, or benchmark/task authoring
How to Apply
Click "Apply" to be taken to the Mercor website. Complete your profile by uploading your resume and confirming your work location. Once verified, you will be matched to opportunities as they arise. Please note that this role cannot support H1B or STEM OPT candidates. Applying through our link supports WFH Bulletin as a referral partner, but you are welcome to apply directly if you prefer.
Responsibilities
• Explore how frontier AI models behave on coding, ML, and analysis tasks and find where they quietly get things wrong
• Turn discovered weaknesses into well-crafted tasks that are hard for models but fair to grade
• Write up findings clearly with evidence and reproducible steps
• Team up with task authors to close loopholes, shortcuts, and grading gaps
• Share insights with researchers and fellow experts to improve the benchmark
Requirements
• MSc or PhD in a STEM field, or equivalent practical experience in a research-heavy domain requiring data analysis and coding
• 1+ years of experience in a research, research-engineering, security, or AI-evaluation role
• Demonstrated ability to identify vulnerabilities, edge cases, or failure modes in LLMs or ML systems (via red teaming, adversarial testing, security research, or rigorous model evaluation)
• Working proficiency in Python and Git, with the ability to script your own probes and analyses
• Strong familiarity with LLM capabilities, limitations, and evaluation techniques
• Perfectionist mindset with high attention to detail, creativity in finding what others missed, strong written communication, and ability to work independently on ambiguous problems
• Ability to engage reliably for ~35 hours per week
Preferred
• Past experience in AI training, model evaluation, or benchmark/task authoring
How to Apply
Click "Apply" to be taken to the Mercor website. Complete your profile by uploading your resume and confirming your work location. Once verified, you will be matched to opportunities as they arise. Please note that this role cannot support H1B or STEM OPT candidates. Applying through our link supports WFH Bulletin as a referral partner, but you are welcome to apply directly if you prefer.


Mercor
LLM Red Team Specialist
LLM Red Team Specialist
Overview
Overview
Cincinnatus LLC is recruiting LLM Red Team Specialists for a leading AI lab to find where frontier models break on complex, multi-step tasks. You will probe models, identify vulnerabilities and edge cases, design challenging tasks around those weaknesses, and document findings. This is a full-time W-2 contingent role (~35 hours/week) with the opportunity to be placed at a leading AI lab.
Cincinnatus LLC is recruiting LLM Red Team Specialists for a leading AI lab to find where frontier models break on complex, multi-step tasks. You will probe models, identify vulnerabilities and edge cases, design challenging tasks around those weaknesses, and document findings. This is a full-time W-2 contingent role (~35 hours/week) with the opportunity to be placed at a leading AI lab.
Cincinnatus LLC is recruiting LLM Red Team Specialists for a leading AI lab to find where frontier models break on complex, multi-step tasks. You will probe models, identify vulnerabilities and edge cases, design challenging tasks around those weaknesses, and document findings. This is a full-time W-2 contingent role (~35 hours/week) with the opportunity to be placed at a leading AI lab.
Get Started
Find Verified Remote Jobs That Fit Your Career Goals
Explore carefully reviewed remote job opportunities from trusted companies worldwide. Discover roles that match your skills, experience and work preferences all in one place.
Newsletter
Get Started
Find Verified Remote Jobs That Fit Your Career Goals
Explore carefully reviewed remote job opportunities from trusted companies worldwide. Discover roles that match your skills, experience and work preferences all in one place.
Newsletter
Get Started
Find Verified Remote Jobs That Fit Your Career Goals
Explore carefully reviewed remote job opportunities from trusted companies worldwide. Discover roles that match your skills, experience and work preferences all in one place.

