Question about different seeds per gpu with DDP #239

HIT-LiuChen · 2024-01-25T08:38:08Z

Line 182 in 35cd455

seed = args.seed + utils.get_rank()

In the issue Why should we set different seed per gpu with DDP, the explanation is that the different seed contributes to the not same data-augmentations on different GPUs. However, I have another question. The different seeds on different GPUs also make different model weight initialization. I dont find the synchronous code like torch.distributed.boardcast(). Is the different initilization helpful in distributed training process? Or, would you provide the synchronous code on model initilization?

The text was updated successfully, but these errors were encountered:

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Question about different seeds per gpu with DDP #239

Question about different seeds per gpu with DDP #239

HIT-LiuChen commented Jan 25, 2024

Question about different seeds per gpu with DDP #239

Question about different seeds per gpu with DDP #239

Comments

HIT-LiuChen commented Jan 25, 2024