首席科学家/GPU Architect/ESL /SOC Architect/超节点/编译器/大模型技术总监/软件栈架构/ 200-300k·15薪
成都 10年以上 硕士 招20人 8月14日更新
收藏
年终奖金 股票期权 绩效奖金 全勤奖 五险一金 发展空间大 公司规模大 扁平管理 管理规范 团队聚餐
avator
段恩华 1天前在线 已认证
高级猎头顾问 · 武汉云猎人力资源有限公司
简历处理快 回复速度快
聊一聊
职位介绍
AI Chip (Top-Tier) Executive Openings Locations: Beijing / Shanghai / Shenzhen / Hong Kong / Chengdu / Xi’an /,China 地点:北京 / 上海 / 深圳 / 香港 / 成都 / 西安,中国 Chief Scientis: — Lead the overall technology roadmap and architecture design across AI algorithms, AI processor architecture, and big-data platforms. Drive key innovations in deep learning, large-scale models, heterogeneous computing, and distributed systems, enabling hardware-software co-design. Spearhead AI system performance optimization and industrialization, building next-generation GPGPU + NPU-integrated intelligent computing platforms. CN: 负责AI算法、AI处理器架构与大数据平台的整体技术规划与架构设计,制定核心研发路线。 主导深度学习、大模型、异构计算与分布式系统等关键技术攻关,推动软硬件协同创新。 XXXAI系统性能优化与产业化落地,打造面向智算时代的GPGPU+NPU融合技术体系。 Chief GPGPU Architect: — Define the overall architecture of the company’s flagship GPGPU, leading instruction set, compute core, and interconnect design to deliver industry-leading performance and scalability. CN: 定义公司旗舰级 GPGPU 的整体架构,主导指令集、计算核心与片上互联的设计,实现业界XXX的性能与可扩展性。 GPGPU Performance Modeling Lead: — Lead architecture-level and system-level performance modeling for next-generation GPGPU chips, driving design trade-offs and guiding compute, memory, and interconnect optimization. CN: 主导下一代 GPGPU 芯片的架构级与系统级性能建模,驱动性能/功耗/面积(PPA)权衡分析,指导计算、存储与互联优化。 CUDA Technical Director: — Design a hardware-software co-optimized CUDA compatibility layer to ensure seamless ecosystem migration and break customer “ecosystem lock-in.” Lead efforts to achieve compatibility and recognition from major global vendors. CN: 设计软硬件协同优化的 CUDA 兼容层,确保生态无缝迁移,打破客户的“生态依赖”;带领团队实现与国际主流厂商生态兼容与认证。 AI SoC Chief Architect: — Lead the design of hybrid GPGPU + NPU architectures. Define core instruction sets and compute allocation logic—the “skeleton” of the chip—to balance inference performance and power efficiency. CN: 主导 GPGPU + NPU 混合架构设计,定义核心指令集与算力分配逻辑,构建芯片的“骨架结构”,在推理性能与能效比间实现XXX平衡。 GPGPU Compute Core Expert / Director Level: — Architect and design GPGPU+ NPU streaming processors and scheduling units to ensure robust general-purpose compute capability. Enable flexible compute for non-AI workloads to meet diverse enterprise needs. CN: 负责 GPGPU + NPU 流处理器及调度单元的架构与设计,保障通用计算能力;支持非 AI 负载的灵活算力,满足多样化企业应用需求。 Director of Large Model Inference: — Lead the overall architecture and performance optimization of large-scale model inference systems across algorithm, system, and hardware layers. Drive large-model (10B–1T+) inference deployment and acceleration on GPU clusters, achieving state-of-the-art stability, latency, and efficiency. Guide the team in kernel-level optimization (CUDA/NCCL/Triton) and cross-layer co-design, covering operator fusion, graph optimization, and memory scheduling to deliver best-in-class inference performance. CN: 负责公司大模型推理系统的整体架构设计与性能优化,统筹算法、系统与硬件的协同加速;主导百亿至万亿参数级模型在GPU集群上的推理部署与加速优化,确保在稳定性、延迟和能效方面达到行业XXX; 带领团队进行算子融合、图优化、显存调度及CUDA/NCCL/Triton内核级优化,实现跨层协同的高性能推理系统。 Super-Node System Architect / Director Level: — Responsible for multi-chip interconnect, memory subsystem (e.g., HBM), and network topology design to unleash per-chip performance and achieve synergistic cluster-level efficiency. CN: 负责多芯片互联架构、存储子系统(如 HBM)与网络拓扑设计,释放单芯片算力,实现集群级协同性能提升。 AI Compiler Expert/Lead: — Lead the design of AI compilers for high-performance processors, optimizing from deep learning frameworks to hardware execution. Drive LLVM/TVM/MLIR innovations and collaborate on hardware-software co-design to unlock breakthrough AI performance. CN: 主导高性能AI处理器编译器的设计与开发,构建从深度学习框架到硬件执行的高效编译链路。推动LLVM / TVM / MLIR 等编译技术创新,并通过软硬件协同设计,实现AI算力的突破性提升。 AI Software Stack Architect: — Design and optimize the full software stack for large-scale AI chips, covering drivers, compilers, runtime, and performance tuning. Lead distributed deployment and orchestration for AI acceleration clusters, drawing on DGX and Cloud Stack best practices. Build a robust software ecosystem for self-developed AI accelerators, enabling large-model training and inference at scale. CN: 负责自研大算力AI芯片的软件栈总体设计与性能优化,涵盖驱动层、编译器与Runtime体系。 主导AI集群分布式部署与调度架构设计,参考DGX与Cloud Stack的系统实践。 构建自研AI加速卡的软件生态,支撑大模型训练与推理的规模化落地。
其他信息
语言要求:普通话
行业要求:人工智能

猎聘温馨提示:

1. 如您发现平台内招聘方存在以下违规行为的,请立即举报
  • · 扣押您的身份证件或者其他证件;
  • · 要求您提供担保人、担保金或者以其他名义向您收取财物( 如培训费、体检费、资料费、置装费、押金等);
  • · 强迫您入股或者向您集资;
  • · 以招聘名义牟取不正当利益;
  • · 发布虚假招聘广告信息;
  • · 工作时长违反劳动法规定;
  • · 存在其他损害您的合法权益的行为。
2. 如您应聘的岗位属于涉外劳务合作/海外岗位的,请务必核实招聘方对外劳务合作资质取得情况,同时注意自身资金安全,防范招聘欺诈。
3. 本平台招聘方不向求职者提供任何收费服务。
查看全部
更新时间:2026-08-15