[EZ] Fix set_torch_num_threads in multi-node. #2164

EugenHotaj · 2024-12-17T14:53:51Z

Current code assumes we only have a single node and sets num_threads incorrectly. This will fail entirely when world_size > num_threads.

Context

What is the purpose of this PR? Is it to

add a new feature
fix a bug
update tests and/or documentation
other (please add here)

Please link to any issues this PR addresses.

Changelog

What are the changes made in this PR?
*

Test plan

Please make sure to do each of the following if applicable to your PR. If you're unsure about any one of these just ask and we will happily help. We also have a contributing page for some guidance on contributing.

run pre-commit hooks and linters (make sure you've first installed via pre-commit install)
add unit tests for any new functionality
update docstrings for any new or updated methods or classes
run unit tests via pytest tests
run recipe tests via pytest tests -m integration_test
manually run any new or modified recipes with sufficient proof of correctness
include relevant commands and any other artifacts in this summary (pastes of loss curves, eval results, etc.)

UX

If your function changed a public API, please add a dummy example of what the user experience will look like when calling it.
Here is a docstring example
and a tutorial example

I did not change any public API
I have added an example to docs or docstrings

Current code assumes we only have a single node and sets num_threads incorrectly. This will fail entirely when world_size > num_threads.

pytorch-bot · 2024-12-17T14:53:54Z

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/torchtune/2164

📄 Preview Python docs built from this PR

Note: Links to docs will display an error until the docs builds have been completed.

✅ No Failures

As of commit 0f90c31 with merge base 9dae7f1 ():
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

facebook-github-bot · 2024-12-17T14:53:57Z

Hi @EugenHotaj!

Thank you for your pull request.

We require contributors to sign our Contributor License Agreement, and yours needs attention.

You currently have a record in our system, but the CLA is no longer valid, and will need to be resubmitted.

Process

In order for us to review and merge your suggested changes, please sign at https://code.facebook.com/cla. If you are contributing on behalf of someone else (eg your employer), the individual CLA may not be sufficient and your employer may need to sign the corporate CLA.

Once the CLA is signed, our tooling will perform checks and validations. Afterwards, the pull request will be tagged with CLA signed. The tagging process may take up to 1 hour after signing. Please give it that time before contacting us about it.

If you have received this in error or have any questions, please contact us at cla@meta.com. Thanks!

facebook-github-bot · 2024-12-17T18:12:02Z

Thank you for signing our Contributor License Agreement. We can now accept your code for this (and any) Meta Open Source project. Thanks!

ebsmothers

Thanks for the fix!

* Llama 3.3 70B (meta-pytorch#2124) * Llama 3.3 readme updates (meta-pytorch#2125) * update configs (meta-pytorch#2107) Co-authored-by: Felipe Mello <felipemello@fb.com> * Reduce logging output for distributed KD (meta-pytorch#2120) * Support Early Exit Loss and/or Layer Dropout (meta-pytorch#1076) Co-authored-by: ebsmothers <ebs@meta.com> * Update checkpointing directory (meta-pytorch#2074) Co-authored-by: Felipe Mello <felipemello@fb.com> Co-authored-by: vancoyendall <vancoykendall@gmail.com> * pass correct arg (meta-pytorch#2127) Co-authored-by: Felipe Mello <felipemello@fb.com> * update configs (meta-pytorch#2128) Co-authored-by: Felipe Mello <felipemello@fb.com> * fix qat_lora_test (meta-pytorch#2131) Co-authored-by: Felipe Mello <felipemello@fb.com> * guard ckpt imports (meta-pytorch#2133) Co-authored-by: Felipe Mello <felipemello@fb.com> * [bug fix] add parents=True (meta-pytorch#2136) Co-authored-by: Felipe Mello <felipemello@fb.com> * [bug fix] re-add model (meta-pytorch#2135) Co-authored-by: Felipe Mello <felipemello@fb.com> * Update save sizes into GiB (meta-pytorch#2143) * [bug fix] remove config download when source is kaggle (meta-pytorch#2144) Co-authored-by: Felipe Mello <felipemello@fb.com> * [fix] remove "with_suffix" (meta-pytorch#2146) Co-authored-by: Felipe Mello <felipemello@fb.com> * DoRA fixes (meta-pytorch#2139) Co-authored-by: Mircea Mironenco <5738815+mirceamironenco@users.noreply.github.com> * [Fix] Llama 3.2 Vision decoder_trainable flag fixed (meta-pytorch#2150) * Small readme, config updates (meta-pytorch#2157) * Using `FormattedCheckpointFiles` in configs (meta-pytorch#2147) * Move ``get_world_size_and_rank`` to utils (meta-pytorch#2155) * Faster intermediate checkpoints with DCP async save in TorchTune (meta-pytorch#2006) Co-authored-by: Saurabh Mishra <msaurabh@fb.com> * torchdata integration - multi-dataset and streaming support (meta-pytorch#1929) * Allow higher version of lm-eval (meta-pytorch#2165) * Using `FormattedCheckpointFiles` in configs... round 2 (meta-pytorch#2167) * [EZ] Fix set_torch_num_threads in multi-node. (meta-pytorch#2164) --------- Co-authored-by: Philip Bontrager <pbontrager@gmail.com> Co-authored-by: ebsmothers <ebs@meta.com> Co-authored-by: Felipe Mello <fmellomascarenhas@gmail.com> Co-authored-by: Felipe Mello <felipemello@fb.com> Co-authored-by: Joe Cummings <jrcummings27@gmail.com> Co-authored-by: Mostafa Elhoushi <m.elhoushi@ieee.org> Co-authored-by: vancoyendall <vancoykendall@gmail.com> Co-authored-by: Mircea Mironenco <5738815+mirceamironenco@users.noreply.github.com> Co-authored-by: salman <salman.mohammadi@outlook.com> Co-authored-by: Saurabh Mishra <msaurabh@meta.com> Co-authored-by: Saurabh Mishra <msaurabh@fb.com> Co-authored-by: Andrew Ho <andrew.kenneth.ho@gmail.com> Co-authored-by: Eugen Hotaj <eugen_hotaj_91@hotmail.com>

[EZ] Fix set_torch_num_threads in multi-node.

0f90c31

Current code assumes we only have a single node and sets num_threads incorrectly. This will fail entirely when world_size > num_threads.

facebook-github-bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Dec 17, 2024

ebsmothers approved these changes Dec 18, 2024

View reviewed changes

ebsmothers merged commit 1372af4 into meta-pytorch:main Dec 18, 2024
17 checks passed

felipemello1 pushed a commit that referenced this pull request Dec 20, 2024

[EZ] Fix set_torch_num_threads in multi-node. (#2164)

8600c49

mori360 pushed a commit to mori360/torchtune that referenced this pull request Dec 20, 2024

[EZ] Fix set_torch_num_threads in multi-node. (meta-pytorch#2164)

2a3ebca

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Uh oh!

[EZ] Fix set_torch_num_threads in multi-node. #2164

[EZ] Fix set_torch_num_threads in multi-node. #2164

Uh oh!

EugenHotaj commented Dec 17, 2024

Uh oh!

pytorch-bot bot commented Dec 17, 2024 •

edited

Loading

Uh oh!

facebook-github-bot commented Dec 17, 2024

Uh oh!

facebook-github-bot commented Dec 17, 2024

Uh oh!

ebsmothers left a comment

Uh oh!

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

3 participants

[EZ] Fix set_torch_num_threads in multi-node. #2164

[EZ] Fix set_torch_num_threads in multi-node. #2164

Uh oh!

Conversation

EugenHotaj commented Dec 17, 2024

Context

Changelog

Test plan

UX

Uh oh!

pytorch-bot bot commented Dec 17, 2024 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/torchtune/2164

✅ No Failures

Uh oh!

facebook-github-bot commented Dec 17, 2024

Process

Uh oh!

facebook-github-bot commented Dec 17, 2024

Uh oh!

ebsmothers left a comment

Choose a reason for hiding this comment

Uh oh!

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

3 participants

pytorch-bot bot commented Dec 17, 2024 •

edited

Loading