You're Invited:Meet the Socket Team at BlackHat and DEF CON in Las Vegas, Aug 4-6.RSVP →

Book a Demo Install Sign in

cloud-tpu-diagnostics

Package Overview

Advanced tools

Install Socket

Detect and block malicious and high-risk dependencies

Install

cloud-tpu-diagnostics

Monitor, debug and profile the jobs running on Cloud TPU.

0.1.5

PyPI

Maintainers: 2

Cloud TPU Diagnostics

This is a comprehensive library to monitor, debug and profile the jobs running on Cloud TPU. To learn about Cloud TPU, refer to the full documentation.

Features

1. Debugging

1.1 Collect Stack Traces

This module will dump the python traces when a fault such as Segmentation fault, Floating-point exception, Illegal operation exception occurs in the program. Additionally, it will also periodically collect stack traces to help debug when a program running on Cloud TPU is stuck or hung somewhere.

Installation

To install the package, run the following command on TPU VM:

pip install cloud-tpu-diagnostics

Usage

To use this package, first import the module:

from cloud_tpu_diagnostics import diagnostic
from cloud_tpu_diagnostics.configuration import debug_configuration
from cloud_tpu_diagnostics.configuration import diagnostic_configuration
from cloud_tpu_diagnostics.configuration import stack_trace_configuration

Then, create configuration object for stack traces. The module will only collect stack traces when collect_stack_trace parameter is set to True. There are following scenarios supported currently:

Scenario 1: Do not collect stack traces on faults

stack_trace_config = stack_trace_configuration.StackTraceConfig(
                      collect_stack_trace=False)

This configuration will prevent you from collecting stack traces in the event of a fault or process hang.

Scenario 2: Collect stack traces on faults and display on console

stack_trace_config = stack_trace_configuration.StackTraceConfig(
                      collect_stack_trace=True,
                      stack_trace_to_cloud=False)

If there is a fault or process hang, this configuration will show the stack traces on the console (stderr).

Scenario 3: Collect stack traces on faults and upload on cloud

stack_trace_config = stack_trace_configuration.StackTraceConfig(
                      collect_stack_trace=True,
                      stack_trace_to_cloud=True)

This configuration will temporary collect stack traces inside /tmp/debugging directory on TPU host if there is a fault or process hang. Additionally, the traces collected in TPU host memory will be uploaded to Google Cloud Logging, which will make it easier to troubleshoot and fix the problems. You can view the traces in Logs Explorer using the following query:

logName="projects/<project_name>/logs/tpu.googleapis.com%2Fruntime_monitor"
jsonPayload.verb="stacktraceanalyzer"

By default, stack traces will be collected every 10 minutes. In order to change the duration between two stack trace collection events, add the following configuration:

stack_trace_config = stack_trace_configuration.StackTraceConfig(
                      collect_stack_trace=True,
                      stack_trace_to_cloud=True,
                      stack_trace_interval_seconds=300)

This configuration will collect the stack traces on cloud after every 5 minutes.

Then, create configuration object for debug.

debug_config = debug_configuration.DebugConfig(
                stack_trace_config=stack_trace_config)

Then, create configuration object for diagnostic.

diagnostic_config = diagnostic_configuration.DiagnosticConfig(
                      debug_config=debug_config)

Finally, call the diagnose() method using with and wrap the statements inside the context manager for which you want to collect the stack traces.

with diagnostic.diagnose(diagnostic_config):
    run_job(...)

FAQs

What is cloud-tpu-diagnostics?

Is cloud-tpu-diagnostics well maintained?

Did you know?

Socket for GitHub automatically highlights issues in each pull request and monitors the health of all your open source dependencies. Discover the contents of your packages and block harmful activity before you install or update your dependencies.

Install

cloud-tpu-diagnostics

Cloud TPU Diagnostics

Features

1. Debugging

1.1 Collect Stack Traces

Installation

Usage

Scenario 1: Do not collect stack traces on faults

Scenario 2: Collect stack traces on faults and display on console

Scenario 3: Collect stack traces on faults and upload on cloud

Related posts

AI + a16z Podcast: Vibe Coding, Security Risks, and the Path to Progress

Toptal’s GitHub Organization Hijacked: 10 Malicious Packages Published