Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code)
Mercor is partnering with leading AI labs on a recent benchmark for scientific computing. You will author original, executable research problems that today's frontier models cannot solve.
Domains - depth required in at least two subdomains (with a coding focus)
- Mathematics — numerical linear algebra, computational mechanics, computational finance
What you'll do
- Source your own material: a published paper, a Kaggle dataset, an open-source repository, or a scenario you design
- Write scientific prompts based on the input
- Build the grading criteria that define a correct answer
- Calibrate against frontier models — a task ships only when strong models fail it more often than they succeed
Required
- PhD in mathematics, applied mathematics, computational mathematics, or a closely related field
- Demonstrated depth in at least two of the following subdomains: numerical linear algebra, computational mechanics, computational finance
- Working proficiency in Python or R for scientific computing
- Comfortable with Git/GitHub and running code in Docker — authoring runs through a pull-request workflow with automated quality checks
Preferred
- Publications in peer-reviewed journals
- Prior scientific software or research engineering experience
1. Upload your resume and application form
2. A 25-minute conversational interview covering your background, experience, and motivations
3. Follow up within a few days with next steps and onboarding
📌 Mathematics PhD - AI Evaluation Expert (Toronto)
🏢 Obsidian
📍 Toronto
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.