עדיין מחפשים עבודה במנועי חיפוש? הגיע הזמן להשתדרג!
במקום לעבור לבד על אלפי מודעות, Jobify מנתחת את קורות החיים שלך ומציגה לך רק משרות שבאמת מתאימות לך.
מעל 80,000 משרות • 4,000 חדשות ביום
חינם. בלי פרסומות. בלי אותיות קטנות.
Minimum qualifications:
- Bachelor's degree or equivalent practical experience.
- 8 years of experience in software development, focusing on building distributed cloud services.
- 5 years of experience in a formal engineering technical leadership role, leading software engineering teams.
- 2 years of experience in LLM training or inference, including performance optimizations, distributed execution, GPU or TPU acceleration, or PyTorch, JAX, or TensorFlow programming.
- Experience integrating generative AI tools or LLM interfaces into workflows.
- Master’s degree or PhD in Engineering, Computer Science, or a related technical field.
- 3 years of experience working in a complex, matrixed organization involving cross-functional, or cross-business projects.
This role is a unique leadership opportunity to technically lead by example an engineering team dedicated to critical Google Cloud Tensor Processing Unit (TPU) software services, which empower Google’s global AI customers. As a Technical Lead (TL), you will lead the architecture and execution of large-scale ML infrastructure, requiring a deep understanding of Large Language Model (LLM) operations from chip level to fleet levels. You will be responsible for solving complex ML infrastructure issues, driving technical strategy, fostering a high-performance team culture, and working directly with customers to ensure successful landings.
Google Cloud accelerates every organization’s ability to digitally transform its business and industry. We deliver enterprise-grade solutions that leverage Google’s cutting-edge technology, and tools that help developers build more sustainably. Customers in more than 200 countries and territories turn to Google Cloud as their trusted partner to enable growth and solve their most critical business problems.
Responsibilities
- Act as the crucial bridge between raw Tensor Processing Unit (TPU) silicon and production-ready machine learning, owning the software integration and operational ecosystem that powers Google's most advanced AI.
- Lead the end-to-end New Product Introduction process—coordinating complex cross-functional launches from initial concept to General Availability—while ensuring the reliability and scalability of a massive fleet of TPU chips.
- Drive foundational engineering efforts, such as developing the TPU runtime API, qualifying the OS images for TPU Virtual Machines (VMs) and Bare Metal instances, and managing fleet-wide reliability through advanced telemetry and automated repair workflows.
- Lead the architecture, technology and the overseeing of implementation of the TPU solutions to production.
- Analyze customers’ issues (working with customers and the field team) and translate these into viable technical solutions.
במקום לעבור לבד על אלפי מודעות, Jobify מנתחת את קורות החיים שלך ומציגה לך רק משרות שבאמת מתאימות לך.
מעל 80,000 משרות • 4,000 חדשות ביום
חינם. בלי פרסומות. בלי אותיות קטנות.