Deep dive
AI Algorithm Development Platform
Capability theme AI productization & platforms
NetEase AI high-performance deep-learning platform: host training jobs and model services, unify multi-framework GPU cluster scheduling, so teams focus on algorithms—not infra.
Verifiable signals
Reached ~80% of internal algorithm teams; 10+ external customers.
Verifiable signals
Versus externally purchased commercial GPU systems, large-scale compute cost savings up to ~60% (per platform materials).
Verifiable signals
Multi-framework distributed training speedup near-linear (per platform performance materials).
Verifiable signals
One-click/fast deploy cuts cold-start cost; HA design makes underlying faults transparent to users.
Problem
Self-built deep-learning stacks were costly: each team stood up its own frameworks, clusters were hard to build, utilization was low, cross-platform work was fragmented, and maintenance was heavy—algorithm teams were stuck on infra with high compute cost and poor scale.
My role
Owned positioning and architecture, product standards, pricing/partnership models, and full lifecycle from inception to go-to-market—driving internal scale adoption and external expansion.
What I owned
- Platform value: high-performance DL compute that hosts training jobs and model services, hiding env setup, tuning, and resource-management complexity.
- Training compute: unified GPU scheduling for TensorFlow / PyTorch / Caffe and distributed training; large-scale sparse distributed compute for CTR/recommendation-style workloads.
- Inference services: standardized GPU/CPU Docker inference; one-click deploy from trained models; pruning/quantization and unified inference tooling as expansion.
- Productization & commercial: standards, usability, HA; pricing/partnerships from positioning/competitors/infra cost for internal adoption and external customers.
What I did not own
- Business algorithm quality itself (CV/speech/recommendation owned by algo teams; platform provides compute and managed engineering).
- Physical datacenter/network build-out (platform consumes HPC network/storage—does not replace infra orgs).
Collaboration
- Algorithm teams: vision, speech, and data-intelligence teams as core users driving framework and job-shape needs.
- Data & storage: collaborate with big-data compute/storage (e.g. HDFS, parallel NAS, and group data-source connectivity).
- Commercial: external expansion and partnership design under infra-cost constraints.
Requirements analysis
The following expands resume scope against public industry methods (OTD/configurable BOM or MLOps). Findings are checkable; unconfirmed internal details are not invented.
1. Pain-driven problem statements
Materials collapse self-build DL pains into three: low utilization, hard env setup, high maintenance. The product need is not “another training script helper”—it is a managed, uniformly scheduled platform.
- Each org standing up frameworks → need unified multi-framework + one-click deploy.
- Fragmented/non-cluster resources → need larger elastic clusters and higher utilization.
- Self-ops → need HA managed service with transparent underlying faults.
2. Service backlog: training / inference / expansion
Existing services emphasize distributed training (TF/Caffe/PyTorch, etc.) and standardized inference Docker; expansion includes large-scale sparse CTR/rec training, compression/quantization, unified inference engines, and heterogeneous-device cost exploration. Requirements split P0 main path vs later expansion.
- P0: multi-framework training hosting + one-click inference deploy + unified GPU scheduling.
- Expand: ultra-large sparse CTR training, compression toolchain, new heterogeneous compute.
3. Architecture needs: storage · scheduling · compute · algo jobs
Tech architecture layers storage (HDFS/NAS, etc.), big-data compute, GPU/CPU HPC networking, job scheduling, and multi-framework training jobs. Product requirements must ensure data can enter, training jobs run stably, and inference can ship.
- Storage diversity is a requirement: parallel NAS, HDFS, and group data-source connectivity.
- Job shapes cover CV / NLP / Speech / CTR-style training jobs.
4. Acceptance & value signals
Acceptance looks at adoption and cost: internal algo-team coverage, external customers; savings vs purchased commercial systems; near-linear distributed speedup; and whether real cases run on the platform (vision/speech/recommendation, etc.).
- Cost: large-scale compute savings ~60% vs purchased commercial systems.
- Adoption: ~80% internal algo teams + 10+ external customers.
- Case domains: computer vision, speech/language, data intelligence/recommendation, etc.
Flows & architecture schematics
Copy/diagrams integrate resume metrics with public capability descriptions from the NetEase AI Deep Learning Platform intro; internal cluster topology and unauthorized business details are omitted.
Self-build pains
Platform path
From NetEase AI Deep Learning Platform intro: replace per-team self-build with a unified platform that productizes training hosting and model services.
Layered from NetEase AI Deep Learning Platform intro: managed unified scheduling for multi-framework train/infer; internal cluster topology omitted.
Approach
- Replace self-build stacks with a managed platform: unified GPU scheduling and multi-framework support to cut env/ops burden.
- Connect train to infer: distributed training jobs + one-click Docker inference for a repeatable delivery path.
- Productize with cost/performance signals: utilization, near-linear speedup, savings vs purchase—and design pricing/partnerships.
Outcomes
- Shipped a high-performance DL compute platform for priority group businesses—training hosting and model services.
- Reached ~80% of internal algorithm teams; expanded to 10+ external customers.
- Established narratable platform advantages on cost/performance (~60% cost-savings claim; near-linear distributed speedup).