HG
Full time
🤖 AI / ML
AI Engineer - LLM/Agentic AI at Hula Global
Listed by DevShelfHub
Bangalore, India
Hybrid
Mid-level · 3–5 yrs
About the Role
Real-time conversational AI requires engineers who can build async multi-agent systems with sub-second latency at scale. Hula Global is hiring an AI Engineer specializing in LLM and Agentic AI in Bangalore.
What you'll do
- Design and implement asynchronous multi-agent orchestration
- Own end-to-end latency from user message to AI response
- Build resilient inference pipelines under load
- Implement request routing and load balancing for AI workloads
- Migrate AI conversation flow from monolith to dedicated services
- Implement WebSocket/streaming infrastructure for real-time chat
- Design circuit breakers and fallback strategies for AI model failures
- Build observability for AI system performance
- Optimize caching strategies for credit data retrieval
What we're looking for
- 3-5 years building production systems handling 10k+ concurrent users
- Proven experience with async/event-driven architectures
- Hands-on experience scaling ML/AI inference in production
- Deep understanding of caching strategies (Redis, in-memory, CDN)
Skills & Technologies
Required
3-5 years building production systems handling 10k+ concurrent users
Proven experience with async/event-driven architectures
Hands-on experience scaling ML/AI inference in production
Deep understanding of caching strategies (Redis, in-memory, CDN)
Experience with message queues and real-time communication protocols
Built systems integrating multiple LLM/AI models in production
Experience with AI model serving frameworks (TensorFlow Serving, Triton)
Understanding of AI inference optimization (batching, caching, quantization)
Benefits & Perks
- Health Insurance
- Flexible Leave Policy
- Learning Budget
- EPF / NPS
Why This Role is Good for Experienced Professionals
- This hands-on role owns the full stack of agentic AI infrastructure from orchestration to real-time delivery.
- Design and implement asynchronous multi-agent orchestration
- Own end-to-end latency from user message to AI response
- Build resilient inference pipelines that gracefully degrade under load
- Implement intelligent request routing and load balancing
- Migrate AI conversation flow from monolith to dedicated services
- Implement WebSocket/streaming infrastructure for real-time chat
- Design circuit breakers and fallback strategies for AI model failures
- Build comprehensive observability for AI system performance
- Optimize credit data retrieval and caching strategies
- Work with multiple LLM/AI models in production environments
About Hula Global
Hula Global is building production AI systems for real-time conversational experiences. This role focuses on LLM/agentic AI infrastructure — orchestration, streaming, inference optimization, and observability for high-concurrency workloads.
- Industry
- Technology
- Company Size
- 500+
- Website
- hulaglobal.com
🚀 Boost your chances
Skip the queue
Contact the recruiter directly at Hula Global and follow up personally — most applicants never do.
Browse Recruiter DatabaseFree trial · 40 contacts · ₹0
Job Details
- Type
- Full time
- Level
- Mid-level
- Experience
- 3–5 years
- Location
- Bangalore, India
- Work Mode
- Hybrid
- Category
- AI / ML
Share this Job
Recruiter Database
Want to stand out? Talk to the recruiter directly.
Most applicants never hear back. Our Recruiter Database gives you direct email access to the people making hiring decisions — so you can follow up personally and actually get noticed.
- Recruiter name, company & direct email address
- Sourced from active job postings across top companies
- Free trial — 40 contacts, no credit card needed