SIGN IN SIGN UP

PR #459: Add Google Cloud ML Diagnostics metrics support and documentation

Imported from GitHub PR https://github.com/AI-Hypercomputer/maxdiffusion/pull/459

Integrate Google Cloud ML Diagnostics SDK (google-cloud-mldiagnostics) into MaxDiffusion to automatically record training, system, and performance metrics.

Key Changes:
- train_utils.py: Added _METRICS_TO_MANAGED mapping table converting metrics names to canonical MetricType enums (loss, learning_rate, gradient_norm, total_weights, step_time, tflops) with automatic pass-through for custom metrics.
- max_utils.py: configured region=None for GCP cluster auto-discovery, and enabled background hardware metric collection (log_system_metrics=True).
- metrics.md: Created comprehensive guide for capturing metrics using google-cloud-mldiagnostics integration in training scripts.

Tested:
- Ran a multi-host distributed training run on TPU v6e cluster. Verified successful metric ingestion in Cloud Logging for predefined, custom, and hardware utilization metrics.
Copybara import of the project:

--
8457cc7325d28aeaec1ef2d3589f6b0549c86766 by Richa Gupta <richaguptaa@google.com>:

Add Google Cloud ML Diagnostics metrics support and developer guide.

Integrate Google Cloud ML Diagnostics SDK (google-cloud-mldiagnostics) into
MaxDiffusion to automatically record training, system, and performance metrics.

Key Changes:
- train_utils.py: Added _METRICS_TO_MANAGED mapping table converting MaxDiffusion
  keys to canonical MetricType enums (loss, learning_rate, gradient_norm,
  total_weights, step_time, tflops) with automatic pass-through for custom metrics.
  Added batch metric logging to write_metrics().
- max_utils.py: Added _clean_config_dict() to sanitize non-JSON serializable
  hyperparameters, configured region=None for GCP cluster auto-discovery, and
  enabled background hardware metric collection (log_system_metrics=True).
- docs/metrics.md: Created comprehensive integration, architecture, and verification
  guide for developers adding new model trainers.

Tested:
- Ran a 500-step multi-host distributed training run on TPU v6e cluster
  richa-maxdiffusion-test (JobSet richa-metrics-test-v11). Verified
  successful metric ingestion in Cloud Logging (ml_diagnostics_metric) for
  predefined, custom, and hardware utilization metrics.

Merging this change closes #459

COPYBARA_INTEGRATE_REVIEW=https://github.com/AI-Hypercomputer/maxdiffusion/pull/459 from richaguptaa17:mldiagnostics-metrics 8457cc7325d28aeaec1ef2d3589f6b0549c86766
PiperOrigin-RevId: 966686006
R
richaguptaa17 committed
670ff95da237a484c690bc6254464dd590c58a7b
Parent: dddc939
Committed by maxdiffusion authors <google-ml-automation@google.com> on 8/21/2026, 10:57:50 PM