PR #459: Add Google Cloud ML Diagnostics metrics support and documentation
Imported from GitHub PR https://github.com/AI-Hypercomputer/maxdiffusion/pull/459 Integrate Google Cloud ML Diagnostics SDK (google-cloud-mldiagnostics) into MaxDiffusion to automatically record training, system, and performance metrics. Key Changes: - train_utils.py: Added _METRICS_TO_MANAGED mapping table converting metrics names to canonical MetricType enums (loss, learning_rate, gradient_norm, total_weights, step_time, tflops) with automatic pass-through for custom metrics. - max_utils.py: configured region=None for GCP cluster auto-discovery, and enabled background hardware metric collection (log_system_metrics=True). - metrics.md: Created comprehensive guide for capturing metrics using google-cloud-mldiagnostics integration in training scripts. Tested: - Ran a multi-host distributed training run on TPU v6e cluster. Verified successful metric ingestion in Cloud Logging for predefined, custom, and hardware utilization metrics. Copybara import of the project: -- 8457cc7325d28aeaec1ef2d3589f6b0549c86766 by Richa Gupta <richaguptaa@google.com>: Add Google Cloud ML Diagnostics metrics support and developer guide. Integrate Google Cloud ML Diagnostics SDK (google-cloud-mldiagnostics) into MaxDiffusion to automatically record training, system, and performance metrics. Key Changes: - train_utils.py: Added _METRICS_TO_MANAGED mapping table converting MaxDiffusion keys to canonical MetricType enums (loss, learning_rate, gradient_norm, total_weights, step_time, tflops) with automatic pass-through for custom metrics. Added batch metric logging to write_metrics(). - max_utils.py: Added _clean_config_dict() to sanitize non-JSON serializable hyperparameters, configured region=None for GCP cluster auto-discovery, and enabled background hardware metric collection (log_system_metrics=True). - docs/metrics.md: Created comprehensive integration, architecture, and verification guide for developers adding new model trainers. Tested: - Ran a 500-step multi-host distributed training run on TPU v6e cluster richa-maxdiffusion-test (JobSet richa-metrics-test-v11). Verified successful metric ingestion in Cloud Logging (ml_diagnostics_metric) for predefined, custom, and hardware utilization metrics. Merging this change closes #459 COPYBARA_INTEGRATE_REVIEW=https://github.com/AI-Hypercomputer/maxdiffusion/pull/459 from richaguptaa17:mldiagnostics-metrics 8457cc7325d28aeaec1ef2d3589f6b0549c86766 PiperOrigin-RevId: 966686006
R
richaguptaa17 committed
670ff95da237a484c690bc6254464dd590c58a7b
Parent: dddc939
Committed by maxdiffusion authors <google-ml-automation@google.com>
on 8/21/2026, 10:57:50 PM