Jobiglo

لا توجد نتائج.

هذه الوظيفة لم تعد متاحة

انتهت صلاحية هذه الوظيفة في 21/09/2026. لم تعد تقبل الطلبات.

RLHF Specialist (Remote)

Odixcity Consulting

Remote
Remote Mid 🇬🇧 English
Python PyTorch JAX TensorFlow PPO LoRA QLoRA Llama 2 Mistral Gemma LabelBox Scale AI Snorkel AWS SageMaker GCP Vertex AI

وصف الوظيفة

About the role

We are looking for an RLHF Specialist to design and operate reinforcement‑learning‑from‑human‑feedback pipelines that improve the safety, factuality, and alignment of large language models. The role is fully remote and collaborates closely with machine‑learning engineers.

Key responsibilities

  • Generate high‑quality preference data by ranking model responses on helpfulness, honesty, and harmlessness.
  • Design multi‑turn prompts to stress‑test model reasoning and safety.
  • Write detailed chain‑of‑thought explanations to train reward models.
  • Collaborate with ML engineers to analyse failure modes and fill data gaps.
  • Develop and iterate annotation strategies, ensuring consistency across a global team.
  • Probe models for biases, hallucinations, and vulnerabilities, documenting findings.
  • Analyse edge cases where reward models behave unexpectedly and suggest data interventions.
  • Create templated instruction sets for large annotation teams and translate RL concepts into repeatable tasks for junior reviewers.
  • Maintain a personal benchmark set and regularly re‑evaluate new model versions.

Required profile

  • Minimum 2 years experience in data annotation, model evaluation, computational linguistics, or AI trust & safety.
  • Strong proficiency in Python and deep‑learning frameworks such as PyTorch, JAX, or TensorFlow.
  • Deep understanding of reinforcement‑learning concepts (PPO, trust‑region methods, reward hacking) applied to language generation.
  • Hands‑on experience fine‑tuning open‑source models (e.g., Llama 2/3, Mistral, Gemma) using LoRA/QLoRA.
  • Experience with annotation platforms (LabelBox, Scale AI, Snorkel) and human‑in‑the‑loop workflows.
  • Ability to diagnose RL policy collapse and adjust hyper‑parameters or reward structures.
  • Familiarity with constitutional AI, self‑alignment techniques, and contributions to open‑source alignment libraries.
  • Experience with cloud platforms such as AWS SageMaker or GCP Vertex AI.

Required skills

  • Python
  • PyTorch
  • JAX
  • TensorFlow
  • PPO
  • LoRA
  • QLoRA
  • Llama 2/3
  • Mistral
  • Gemma
  • LabelBox
  • Scale AI
  • Snorkel
  • AWS SageMaker
  • GCP Vertex AI

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec Odixcity Consulting.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.

لماذا تبلغ عن هذا العرض؟

شكراً لإبلاغك. سنراجع هذا العرض.
💬 راسلنا على تيليجرام الدردشة عبر واتساب

منشور منذ شهرين

48 مشاهدات · 0 مهتم

عزز فرصك

حمّل سيرتك الذاتية وسنقترح عليك الوظائف التي تناسب ملفك.

جاري تحليل سيرتك الذاتية...

Odixcity Consulting