25 Aug
|
Jobgether
|
Canada
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an AI Safety Expert — English & Odia based in Canada .
This is a remote opportunity for experts who want to help make conversational AI systems safer, more reliable, and more resilient. You will take part in adversarial testing designed to uncover vulnerabilities that automated evaluations may overlook. Your work will involve probing AI models and agents for jailbreaks, prompt injection, misuse scenarios, bias, and multi-turn manipulation.
You will transform your findings into structured annotations, attack cases, datasets, and actionable reports. The role offers exposure to cutting-edge AI safety projects and the opportunity to contribute directly to the development of more trustworthy AI systems. Higher-sensitivity projects are optional and supported by transparent guidelines and wellness resources.
n
Accountabilities:
- Red-team conversational AI models and agents by testing jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulation.
- Identify vulnerabilities, failure modes, systemic risks, and safety gaps that automated testing may not detect.
- Generate high-quality human data by annotating model failures, classifying vulnerabilities, and documenting adversarial examples.
- Apply established taxonomies, benchmarks, testing frameworks, and playbooks to ensure consistent and reproducible evaluations.
- Create structured attack cases, datasets, evaluation artifacts, and reports that can be used to improve AI safety and model performance.
- Clearly communicate technical and behavioral risks to both technical and non-technical stakeholders.
- Adapt testing approaches across different AI safety projects, models, use cases, and customer requirements.
- Contribute specialized expertise in areas such as adversarial machine learning, cybersecurity, socio-technical risk, abuse analysis, psychology, creative writing, or unconventional adversarial testing.
- Work independently in a fully remote environment while maintaining high standards for quality, documentation, and consistency.
- Follow established safety guidelines when reviewing potentially sensitive content, with participation in higher-sensitivity projects remaining optional.
Requirements
- Native-level fluency in English and Odia is required.
- Prior experience with AI red teaming, adversarial AI, cybersecurity, socio-technical probing, AI evaluation, or another closely related field.
- Strong curiosity and an adversarial mindset, with a natural ability to identify ways systems could be manipulated or pushed beyond their intended behavior.
- Ability to approach testing systematically using frameworks, benchmarks, taxonomies, and structured methodologies rather than relying solely on ad hoc experimentation.
- Strong written and verbal communication skills, with the ability to explain vulnerabilities, risks, and findings clearly.
- Excellent analytical, critical-thinking, problem-solving, and documentation skills.
- Ability to work independently, manage priorities, and adapt quickly across different projects and requirements.
- Experience with adversarial machine learning, jailbreak datasets, prompt injection, RLHF/DPO attacks, or model extraction is an advantage.
- Cybersecurity experience such as penetration testing, exploit development, reverse engineering, or security research is a plus.
- Experience with harassment or misinformation analysis, abuse testing, conversational AI evaluation, psychology, acting, or creative writing is also valuable.
- Comfort working with potentially sensitive AI-generated content while following established safety procedures and guidelines.
Benefits
- Fully remote work with the flexibility to complete projects on your own schedule.
- Independent contractor engagement with weekly payments through Stripe or Wise based on services rendered.
- Competitive compensation, with project rates determined according to the applicable expertise and assignment.
- Opportunity to work on cutting-edge AI safety and human-data projects.
- Direct contribution to making AI systems more robust, safe, trustworthy, and resilient.
- Exposure to emerging AI evaluation and red-teaming methodologies.
- Opportunities to collaborate with AI researchers and contribute expertise to frontier AI development.
- Projects may be extended, shortened, or concluded early depending on project needs and performance.
- Higher-sensitivity projects are optional, with topics communicated clearly in advance and supported by established guidelines and wellness resources.
- No access to confidential or proprietary information from other employers, clients, or institutions is required.
- Referral opportunities may provide additional earnings where applicable.
nHow Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Why Apply Through Jobgether?
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
📌 AI Safety Experts — English & Odia (Canada)
🏢 Jobgether
📍 Canada