Site Reliability Engineer - ARK Large Model Platform (Singapore)
About the Team The Applied Machine Learning (AML) - Enterprise team provides machine learning platform products on VolcanoEngine with cloud native resource scheduling system which intelligently orchestrates different tasks and jobs with minimised costs of every experiment and maximised resource utilisation, rich modelling tools including customised machine learning tasks and web IDE, and multi-framework high performance model inference services. In 2021, through VolcanoEngine, we released this machine learning infrastructure to the public, to provide more enterprises with reduced costs of computation power, lower barriers to machine learning engineering and deeper developments in AI capabilities. Responsibilities Responsible for Ark Large Model Platform development on Volcano Engine, researching systematic solutions on large model solution implementations and applications in various industries, striving to reduce the IT cost of large model applications, meeting the users' ever-growing demand for intelligent interaction and improving the lifestyle and communications of users in the future world. - Manage and oversee the stability of both control and data aspects of large-scale model systems through effective DevOps practices. - Develop and enhance observability systems for monitoring the stability of large model systems, ensuring high reliability and performance. - Handle super large-scale cluster management and ensure efficient operation and maintenance of large model systems.