← Back to Products

Deep dive

AI Algorithm Development Platform

Capability theme AI productization & platforms

NetEase AI high-performance deep-learning platform: host training jobs and model services, unify multi-framework GPU cluster scheduling, so teams focus on algorithms—not infra.

Verifiable signals

Reached ~80% of internal algorithm teams; 10+ external customers.

Verifiable signals

Versus externally purchased commercial GPU systems, large-scale compute cost savings up to ~60% (per platform materials).

Verifiable signals

Multi-framework distributed training speedup near-linear (per platform performance materials).

Verifiable signals

One-click/fast deploy cuts cold-start cost; HA design makes underlying faults transparent to users.

Problem

Self-built deep-learning stacks were costly: each team stood up its own frameworks, clusters were hard to build, utilization was low, cross-platform work was fragmented, and maintenance was heavy—algorithm teams were stuck on infra with high compute cost and poor scale.

My role

Owned positioning and architecture, product standards, pricing/partnership models, and full lifecycle from inception to go-to-market—driving internal scale adoption and external expansion.

What I owned

  • Platform value: high-performance DL compute that hosts training jobs and model services, hiding env setup, tuning, and resource-management complexity.
  • Training compute: unified GPU scheduling for TensorFlow / PyTorch / Caffe and distributed training; large-scale sparse distributed compute for CTR/recommendation-style workloads.
  • Inference services: standardized GPU/CPU Docker inference; one-click deploy from trained models; pruning/quantization and unified inference tooling as expansion.
  • Productization & commercial: standards, usability, HA; pricing/partnerships from positioning/competitors/infra cost for internal adoption and external customers.

What I did not own

  • Business algorithm quality itself (CV/speech/recommendation owned by algo teams; platform provides compute and managed engineering).
  • Physical datacenter/network build-out (platform consumes HPC network/storage—does not replace infra orgs).

Collaboration

  • Algorithm teams: vision, speech, and data-intelligence teams as core users driving framework and job-shape needs.
  • Data & storage: collaborate with big-data compute/storage (e.g. HDFS, parallel NAS, and group data-source connectivity).
  • Commercial: external expansion and partnership design under infra-cost constraints.

Requirements analysis

The following expands resume scope against public industry methods (OTD/configurable BOM or MLOps). Findings are checkable; unconfirmed internal details are not invented.

1. Pain-driven problem statements

Materials collapse self-build DL pains into three: low utilization, hard env setup, high maintenance. The product need is not “another training script helper”—it is a managed, uniformly scheduled platform.

  • Each org standing up frameworks → need unified multi-framework + one-click deploy.
  • Fragmented/non-cluster resources → need larger elastic clusters and higher utilization.
  • Self-ops → need HA managed service with transparent underlying faults.

2. Service backlog: training / inference / expansion

Existing services emphasize distributed training (TF/Caffe/PyTorch, etc.) and standardized inference Docker; expansion includes large-scale sparse CTR/rec training, compression/quantization, unified inference engines, and heterogeneous-device cost exploration. Requirements split P0 main path vs later expansion.

  • P0: multi-framework training hosting + one-click inference deploy + unified GPU scheduling.
  • Expand: ultra-large sparse CTR training, compression toolchain, new heterogeneous compute.

3. Architecture needs: storage · scheduling · compute · algo jobs

Tech architecture layers storage (HDFS/NAS, etc.), big-data compute, GPU/CPU HPC networking, job scheduling, and multi-framework training jobs. Product requirements must ensure data can enter, training jobs run stably, and inference can ship.

  • Storage diversity is a requirement: parallel NAS, HDFS, and group data-source connectivity.
  • Job shapes cover CV / NLP / Speech / CTR-style training jobs.

4. Acceptance & value signals

Acceptance looks at adoption and cost: internal algo-team coverage, external customers; savings vs purchased commercial systems; near-linear distributed speedup; and whether real cases run on the platform (vision/speech/recommendation, etc.).

  • Cost: large-scale compute savings ~60% vs purchased commercial systems.
  • Adoption: ~80% internal algo teams + 10+ external customers.
  • Case domains: computer vision, speech/language, data intelligence/recommendation, etc.

Flows & architecture schematics

Copy/diagrams integrate resume metrics with public capability descriptions from the NetEase AI Deep Learning Platform intro; internal cluster topology and unauthorized business details are omitted.

Flow: self-build pain → managed platform path (desensitized)

Self-build pains

Hard env setup Low utilization High maintenance

Platform path

Managed scheduling Multi-framework distributed train One-click inference Adopt / expand

From NetEase AI Deep Learning Platform intro: replace per-team self-build with a unified platform that productizes training hosting and model services.

Architecture: storage · scheduling · multi-framework train · inference (desensitized)
Storage & data HDFS / NAS, etc. Business data ingest Unified scheduling GPU cluster scheduling HA managed hosting Training compute TF / PyTorch / Caffe … Distributed training jobs Inference services GPU/CPU Docker One-click deploy Product value Teams focus on algorithms · lower self-build/ops cost · higher utilization Vs purchased commercial systems: ~60% large-scale compute cost savings (per materials)

Layered from NetEase AI Deep Learning Platform intro: managed unified scheduling for multi-framework train/infer; internal cluster topology omitted.

Approach

  • Replace self-build stacks with a managed platform: unified GPU scheduling and multi-framework support to cut env/ops burden.
  • Connect train to infer: distributed training jobs + one-click Docker inference for a repeatable delivery path.
  • Productize with cost/performance signals: utilization, near-linear speedup, savings vs purchase—and design pricing/partnerships.

Outcomes

  • Shipped a high-performance DL compute platform for priority group businesses—training hosting and model services.
  • Reached ~80% of internal algorithm teams; expanded to 10+ external customers.
  • Established narratable platform advantages on cost/performance (~60% cost-savings claim; near-linear distributed speedup).

Appendix & links