DS DevShelfHub Projects · AI tools
HG
Full time 🤖 AI / ML

AI Engineer - LLM/Agentic AI at Hula Global

Listed by DevShelfHub

Bangalore, India
Hybrid
Mid-level · 3–5 yrs

About the Role

Real-time conversational AI requires engineers who can build async multi-agent systems with sub-second latency at scale. Hula Global is hiring an AI Engineer specializing in LLM and Agentic AI in Bangalore.

What you'll do

  • Design and implement asynchronous multi-agent orchestration
  • Own end-to-end latency from user message to AI response
  • Build resilient inference pipelines under load
  • Implement request routing and load balancing for AI workloads
  • Migrate AI conversation flow from monolith to dedicated services
  • Implement WebSocket/streaming infrastructure for real-time chat
  • Design circuit breakers and fallback strategies for AI model failures
  • Build observability for AI system performance
  • Optimize caching strategies for credit data retrieval

What we're looking for

  • 3-5 years building production systems handling 10k+ concurrent users
  • Proven experience with async/event-driven architectures
  • Hands-on experience scaling ML/AI inference in production
  • Deep understanding of caching strategies (Redis, in-memory, CDN)

Skills & Technologies

Required

3-5 years building production systems handling 10k+ concurrent users Proven experience with async/event-driven architectures Hands-on experience scaling ML/AI inference in production Deep understanding of caching strategies (Redis, in-memory, CDN) Experience with message queues and real-time communication protocols Built systems integrating multiple LLM/AI models in production Experience with AI model serving frameworks (TensorFlow Serving, Triton) Understanding of AI inference optimization (batching, caching, quantization)

Benefits & Perks

  • Health Insurance
  • Flexible Leave Policy
  • Learning Budget
  • EPF / NPS

Why This Role is Good for Experienced Professionals

  • This hands-on role owns the full stack of agentic AI infrastructure from orchestration to real-time delivery.
  • Design and implement asynchronous multi-agent orchestration
  • Own end-to-end latency from user message to AI response
  • Build resilient inference pipelines that gracefully degrade under load
  • Implement intelligent request routing and load balancing
  • Migrate AI conversation flow from monolith to dedicated services
  • Implement WebSocket/streaming infrastructure for real-time chat
  • Design circuit breakers and fallback strategies for AI model failures
  • Build comprehensive observability for AI system performance
  • Optimize credit data retrieval and caching strategies
  • Work with multiple LLM/AI models in production environments

About Hula Global

Hula Global is building production AI systems for real-time conversational experiences. This role focuses on LLM/agentic AI infrastructure — orchestration, streaming, inference optimization, and observability for high-concurrency workloads.
Industry
Technology
Company Size
500+
Website
hulaglobal.com
🚀 Boost your chances

Skip the queue

Contact the recruiter directly at Hula Global and follow up personally — most applicants never do.

Browse Recruiter Database

Free trial · 40 contacts · ₹0

Job Details

Type
Full time
Level
Mid-level
Experience
3–5 years
Location
Bangalore, India
Work Mode
Hybrid
Category
AI / ML

Share this Job

Recruiter Database

Want to stand out? Talk to the recruiter directly.

Most applicants never hear back. Our Recruiter Database gives you direct email access to the people making hiring decisions — so you can follow up personally and actually get noticed.

  • Recruiter name, company & direct email address
  • Sourced from active job postings across top companies
  • Free trial — 40 contacts, no credit card needed

₹0

to get started

Get Free Trial