Hi Yudhiesh. Thanks so much! I've yet to run any benchmarking (because it would certainly be costly!), but it should be beneficial for large scale multi-node multi-gpu training jobs, especially ones that require 100's of nodes, where inter-node connectivity is the bottleneck.
We are mainly interested in NCCL+EFA, which is only compatible with H100, A100, and V100 GPUs. But I think others have used EFA for mpi runs.