Responsibilities:
- ML System Development: Design, develop, and maintain scalable and efficient machine learning systems, including writing ML services and APIs.
- Model Deployment: Implement and manage the deployment of machine learning models, including transformer based LLMs, into production environments, ensuring reliability and scalability.
- Infrastructure Management: Collaborate with infrastructure teams to optimize and manage the underlying systems supporting machine learning workflows.
- Data Pipeline Creation: Create robust and efficient data pipelines for collecting, processing, and preparing datasets for machine learning models.
- Collaboration: Work closely with data scientists, researchers, and cross-functional teams to integrate ML solutions into existing software infrastructure.
- Performance Optimization: Continuously optimize and improve the performance of machine learning algorithms and systems.
- Documentation: Develop and maintain documentation for machine learning systems, APIs, and data pipelines to ensure clarity and ease of use for team members.
Our ideal candidates would:
- 3+ years of experience including working on designing multi-component systems
- Strong grasp of one high-level language like Python.
- General awareness of SQL and database design concepts
- Solid understanding of testing fundamentals
- Strong communication skills
- should have prior experience in managing and executing technology products.
- Decent understanding of various Gen AI based ML approaches
- Experience in building agentic architectures using langgraph or similar libraries
Bonus:
- Prior experience working with high-volume, always-available web-applications
- Experience working with cloud
- Knowledge of cloud platforms such as AWS, GCP, or Azure.
- Experience with deploying small and big open source LLMs in production environments using containerization tools like Docker
- Experience in Distributed systems.
- Experience working with Start-up is a plus point.