Transformers begins FSDP2 and expert parallelism work
The project is adding a two-dimensional device mesh so fully sharded data parallel can run with tensor and expert parallelism.
By tensorThe project is adding a two-dimensional device mesh so fully sharded data parallel can run with tensor and expert parallelism.
By tensorMark Shannon’s plan would require explicit sharing of objects across threads, building on free-threading work with runtime checks and freezing.
By rvalue