Would you like to tell me the throughput during your open-dcoder-0.5B pre-training?
In my open-dcoder-0.5B experiment, 10B tokens in 4H800-80G finish in 105.5h, so the throughput is (10/105.5/424)=0.569 B tokens/per GPU/per day.
This throughput differs significantly from that of the traditional autoregressive model Megatron framework. Do you know the specific reason?
@pengzhangzhi
Would you like to tell me the throughput during your open-dcoder-0.5B pre-training?
In my open-dcoder-0.5B experiment, 10B tokens in 4H800-80G finish in 105.5h, so the throughput is (10/105.5/424)=0.569 B tokens/per GPU/per day.
This throughput differs significantly from that of the traditional autoregressive model Megatron framework. Do you know the specific reason?
@pengzhangzhi