gke-ai-troubleshooting-tpu-mxla-hang
Diagnose GKE Cloud TPU multi-slice training hangs (Megascale `HANG_DETECTED` logs and hang monitored events) using the ML Diagnostics `Megascale XLA (MXLA) Hang Analyzer` (`gcloud alpha mldiagnostics monitored-events`) and 1-minute Cloud Monitoring multi-slice latency metrics (`kubernetes.io/container/multislice/*`). Distinguishes XLA compiler/HLO launch divergence and host data-input stalls from TPU chip, SparseCore, ICI, or network fabric faults. Use when multi-slice TPU training jobs freeze without progressing steps, emit `HANG_DETECTED`, or stall in collective operations. Don't use for gradual step-time throughput drops without hangs (use gke-ai-troubleshooting-tpu-performance-degradation) or pod preemption/eviction restarts (use gke-ai-troubleshooting-jobset-interruption).
Unsigned, install at your own risk
UnverifiedThis skill has no cryptographic signature attached. We can't verify the contents match what the publisher intended.
Install this skill
Run this command in your terminal. No account required — it auto-detects your AI tool and installs the skill file.
npx @skills-hub-ai/cli install google-cloud-skills-gke-ai-troubleshooting-tpu-mxla-hangSetup by platform
Install
One-click setup for your editorRun in your project root
npx @skills-hub-ai/cli install google-cloud-skills-gke-ai-troubleshooting-tpu-mxla-hang --target claude-codeInstructions
Security
Reviews (0)
Related skills
Browse all →More from Google Cloud Skills
View source →More Build skills
Browse category →Frequently asked questions about gke-ai-troubleshooting-tpu-mxla-hang
What does the gke-ai-troubleshooting-tpu-mxla-hang skill do?
Diagnose GKE Cloud TPU multi-slice training hangs (Megascale `HANG_DETECTED` logs and hang monitored events) using the ML Diagnostics `Megascale XLA (MXLA) Hang Analyzer` (`gcloud alpha mldiagnostics monitored-events`) and 1-minute Cloud Monitoring multi-slice latency metrics (`kubernetes.io/container/multislice/*`). Distinguishes XLA compiler/HLO launch divergence and host data-input stalls from TPU chip, SparseCore, ICI, or network fabric faults. Use when multi-slice TPU training jobs freeze without progressing steps, emit `HANG_DETECTED`, or stall in collective operations. Don't use for gradual step-time throughput drops without hangs (use gke-ai-troubleshooting-tpu-performance-degradation) or pod preemption/eviction restarts (use gke-ai-troubleshooting-jobset-interruption). It's a reusable SKILL.md instruction set that loads into your AI coding assistant on demand, no prompt engineering, no copy-pasting every session.
How do I install the gke-ai-troubleshooting-tpu-mxla-hang skill?
Run `npx @skills-hub-ai/cli install google-cloud-skills-gke-ai-troubleshooting-tpu-mxla-hang` from your terminal. The CLI writes the SKILL.md to the correct location for your AI tool (e.g. ~/.claude/skills/google-cloud-skills-gke-ai-troubleshooting-tpu-mxla-hang/ for Claude Code or ~/.cursor/skills/ for Cursor with --target cursor) and adds it to your project's .skills.json lockfile.
Which AI tools does gke-ai-troubleshooting-tpu-mxla-hang work with?
gke-ai-troubleshooting-tpu-mxla-hang runs in Claude Code. It follows the open Agent Skills standard (SKILL.md), so the same skill works in every supported tool without modification.
Is the gke-ai-troubleshooting-tpu-mxla-hang skill free?
Yes. Every skill on skills-hub.ai is free and open-source. There are no premium tiers, paywalls, or usage limits. You only pay for whatever AI assistant you're already using.
How do I use gke-ai-troubleshooting-tpu-mxla-hang after installing it?
In Claude Code, type `/google-cloud-skills-gke-ai-troubleshooting-tpu-mxla-hang` (or whatever slash command the skill registers) and the AI follows the skill's instructions immediately. You can also reference it by name in natural language, your AI loads the skill into context when relevant.
Can I share the gke-ai-troubleshooting-tpu-mxla-hang skill with my team?
Yes. Commit your project's .skills.json lockfile and teammates run `npx @skills-hub-ai/cli install` (no args) to install every skill at the exact version you pinned. Organization-scoped installs work via skills-hub.ai organizations.