HPC Engineer 30-60k·15薪
上海 3-5年 本科 招1人 8月16日更新
收藏
avator
肖先生 4小时前在线 已认证
合伙人 · 上海默锐企业管理咨询有限公司
简历处理快 回复速度快
聊一聊
职位介绍
岗位职责 / Responsibilities 我们正在寻找一位 HPC / GPU 集群工程师,协助设计、运营并持续优化公司用于分布式模型训练及高性能计算的大规模 GPU 计算环境。您将端到端负责集群的性能与稳定性——从 GPU 硬件与互联网络,到存储层、调度系统、监控平台,以及保障数百块加速器高效运行的全套工具链。 We are looking for an HPC / GPU Cluster Engineer to help design, operate, and continuously optimize our large-scale GPU compute environment used for distributed model training and high-performance workloads. You will own the performance and reliability of the cluster end to end — from the GPUs and interconnect fabric up through the storage layer, scheduler, monitoring, and the tooling that keeps hundreds of accelerators running efficiently. 任职要求 / Qualifications 具备生产环境下大规模 GPU 或 HPC 集群的运营经验,能够系统性地识别并解决性能瓶颈。 Experience operating large GPU or HPC clusters in production, identifying performance bottlenecks and resolving them systematically. 具备 RDMA 编程及底层调试的实战经验,包括 RDMA verbs / libibverbs、内存注册、队列对、完成队列、RDMA 读写、发送/接收操作及性能基准测试。 Hands-on experience with RDMA programming and low-level RDMA debugging, including RDMA verbs / libibverbs, memory registration, queue pairs, completion queues, RDMA read/write, send/recv, and performance benchmarking. 熟悉分布式训练的核心底层技术,包括 InfiniBand 和/或 RoCEv2、GPUDirect RDMA、GPUDirect Storage、NCCL、CUDA 驱动及 OFED。 Practical experience with the technologies underpinning distributed training, including InfiniBand and/or RoCEv2, GPUDirect RDMA, GPUDirect Storage, NCCL, CUDA drivers, and OFED. 具备 Slurm 等调度系统的使用经验,包括安装配置、队列/分区管理、作业故障排查及监控。 Experience with workload schedulers such as Slurm, including setup, configuration, queue/partition management, job troubleshooting, and monitoring. 具备高性能共享存储或并行/分布式文件系统的搭建与管理经验,如 Lustre、BeeGFS、WEKA、VAST、DDN/ExaScaler 等。 Experience setting up and managing high-performance shared storage or parallel/distributed filesystems such as Lustre, BeeGFS, WEKA, VAST, DDN/ExaScaler, or similar systems. 熟练掌握 Python、Bash,优先具备 C/C++ 能力,用于集群自动化、诊断、基准测试及监控开发。 Solid scripting/programming ability in Python, Bash, and preferably C/C++, for cluster automation, diagnostics, benchmarking, and monitoring. 熟悉 HPC 网络设计,包括阻塞与非阻塞网络架构、InfiniBand、高性能以太网、网络拓扑、拥塞控制及端到端带宽/延迟故障排查。 Familiarity with HPC networking, including blocking vs. non-blocking fabric design, InfiniBand, high-performance Ethernet, topology, congestion, and end-to-end bandwidth/latency troubleshooting. 扎实的 Linux 系统管理能力,包括网络配置、文件系统管理、内核/驱动问题处理、进程与资源管理及性能调试。Strong Linux system administration skills, including networking, filesystems, kernel/driver issues, process/resource management, and performance debugging.
其他信息
语言要求:英语
行业要求:基金/证券/期货

猎聘温馨提示:

1. 如您发现平台内招聘方存在以下违规行为的,请立即举报
  • · 扣押您的身份证件或者其他证件;
  • · 要求您提供担保人、担保金或者以其他名义向您收取财物( 如培训费、体检费、资料费、置装费、押金等);
  • · 强迫您入股或者向您集资;
  • · 以招聘名义牟取不正当利益;
  • · 发布虚假招聘广告信息;
  • · 工作时长违反劳动法规定;
  • · 存在其他损害您的合法权益的行为。
2. 如您应聘的岗位属于涉外劳务合作/海外岗位的,请务必核实招聘方对外劳务合作资质取得情况,同时注意自身资金安全,防范招聘欺诈。
3. 本平台招聘方不向求职者提供任何收费服务。
查看全部

猜你喜欢

朱先生
猎头顾问
王先生
资深顾问(SC)
IDC运维经理
上海-徐汇区
25-45k
某东莞科技推广服务公司
科技推广服务
李女士
猎头顾问/助理
系统运维岗
上海
30-60k·18薪
某上海保险上市公司
保险 已上市 100-499人
陈先生
顾问(C)
系统运维管理岗
上海
29-40k
某西安基金/证券/期货上市公司
基金/证券/期货 已上市 2000-5000人
邹先生
猎头
Data Center Engineer
上海-浦东新区
50-60k·15薪
某上海基金证券公司
基金/证券/期货 融资未公开 100-499人
丁女士
Manager
存储工程师
上海
35-65k·17薪
某上海基金/证券/期货公司
基金/证券/期货 融资未公开 100-499人
刘女士
sc
安全运维工程师 外企
上海
30-60k·15薪
某知名公司
批发/零售 融资未公开 10000人以上
谷先生
猎头
服务器运维工程师
上海
35-65k·18薪
某上海基金/证券/期货公司
基金/证券/期货 融资未公开 100-499人
刘女士
猎头
Application Support-Payment
上海-浦东新区
30-45k
著名德资银行
银行 不需要融资 10000人以上
施女士
Senior Consultant
运维工程师
上海
20-35k·16薪
某国内大型基金/证券/期货公司
基金/证券/期货 战略融资 5000-10000人
苑女士
顾问(C)
田女士
招聘经理
1 2 3 4
更新时间:2026-08-16