Senior Software Engineer - Model Training & AI Evals at Chegg
Listed by DevShelfHub
About the Role
Chegg is hiring a Senior Engineer for its AI team at the intersection of evaluation science, post-training, and foundation model development (R5340, Remote India). You will own end-to-end eval/benchmarking infrastructure and contribute hands-on to post-training pipelines for industry-specific vertical foundation models.
This role targets candidates with LLM-lab pedigree who have shipped models—not only called APIs—and can design task-level evals, synthetic data flywheels, comparative frontier-model benchmarks, and alignment training runs (SFT, RLHF, RLAIF, DPO/PPO).
What you'll do
- Design/own task-level evaluation frameworks for LLM agents and base models
- Build comparative benchmarking pipelines across frontier models with structured failure analysis
- Produce capability gap reports and track version regressions
- Develop domain-specific benchmarks and IAA pipelines
- Drive synthetic data generation to close capability gaps; validate via auto-eval, human review, and performance lift
- Build automated regression suites in CI/CD for fine-tuning/model updates
- Partner with product/curriculum/research on post-training and data-flywheel priorities
- Lead/contribute to SFT, RLHF, RLAIF, and DPO runs from dataset design through eval-gated release
- Curate instruction-tuning/preference datasets; define quality/rejection/dedup pipelines
- Implement alignment techniques (reward modeling, PRMs, constitutional/RLAIF)
- Run ablations/experiments attributing model behavior changes
What we're looking for
- Required
- 5+ years ML/AI engineering; 2–3+ years focused on LLMs
- Direct hands-on experience at an LLM lab, AI research org, or equivalent frontier team (shipped models)
- Familiarity with full model lifecycle: pre-training data, post-training alignment, eval, production deployment
Skills & Technologies
Required
Benefits & Perks
- Health Insurance
- Flexible Leave Policy
- Learning Budget
- EPF / NPS
Why This Role is Good for Experienced Professionals
- This senior AI role owns the eval feedback loop that drives model improvement—rare depth beyond API integration work.
- Own end-to-end eval and benchmarking infrastructure
- Hands-on post-training for vertical foundation models
- Remote India flexibility
- Work across frontier-model comparative analysis
- Domain-specific benchmark design (e.g., STEM/legal/finance/healthcare)
- Synthetic data strategies tied to capability gaps
- CI/CD-integrated regression suites for model quality
- Partnership with product, curriculum, and research teams
- Deep alignment-method exposure (reward models, PRMs, RLAIF)
- High bar for lab-pedigree LLM experience
- Strong SE fundamentals expected alongside ML depth
- Opportunity to translate eval signals into training decisions
About Chegg
- Industry
- Technology
- Company Size
- 500+
- Website
- chegg.com
Skip the queue
Contact the recruiter directly at Chegg and follow up personally — most applicants never do.
Browse Recruiter DatabaseFree trial · 40 contacts · ₹0
Job Details
- Type
- Full time
- Level
- Senior
- Experience
- 5–9 years
- Location
- Remote India
- Work Mode
- WFH
- Category
- AI / ML
Share this Job
Recruiter Database
Want to stand out? Talk to the recruiter directly.
Most applicants never hear back. Our Recruiter Database gives you direct email access to the people making hiring decisions — so you can follow up personally and actually get noticed.
- Recruiter name, company & direct email address
- Sourced from active job postings across top companies
- Free trial — 40 contacts, no credit card needed