Uncategorized

ByteDance Working on Gigantic AI Model to Take on Anthropic’s Best

ByteDance is training a massive new artificial intelligence model with as many as 10 trillion parameters as the Chinese technology company steps up efforts to compete with leading US AI developers.

The model could approach or even exceed the size of Anthropic’s cutting-edge Mythos system, the Financial Times reported on Friday, citing people familiar with the project.

Still in Early Stages

ByteDance’s new model is still at an early stage of training and could contain as many as 10 trillion parameters, according to the report.

Anthropic does not publicly disclose the size of its latest models, but industry estimates cited in the report put Mythos at around 8 trillion parameters. By comparison, Moonshot AI’s Kimi K3 has 2.8 trillion parameters, while Meituan’s LongCat-2.0 and DeepSeek’s V4-Pro are estimated at around 1.6 trillion parameters each.

However, a higher parameter count does not automatically mean a model will perform better. The figure mainly indicates the scale of the model, while performance also depends on areas such as training methods, data and architecture.

The ByteDance model is currently in the pre-training stage. This part of model development typically takes between three and six months before developers move to fine-tuning and other stages ahead of a possible release.

ByteDance Expands Its AI Development

The project forms part of ByteDance’s wider push into advanced AI. Its Seed team focuses on areas including model pre-training, post-training, inference, memory, learning and interpretability. ByteDance also maintains dedicated infrastructure teams working on distributed training and high-performance inference for foundation models.

According to the Financial Times, ByteDance has spent heavily on AI over the past three years, expanding its data centres and hiring researchers. Its Seed model-development team has around 2,000 members in China and overseas and is led by former Google DeepMind scientist Wu Yonghui.

The company has also expanded its Volcano Engine cloud business, which provides AI services to enterprise customers, and has ambitions to develop its own AI chips.

ByteDance Avoids Model Distillation

ByteDance has reportedly followed a more independent approach to developing its models for more than a year instead of relying on model distillation from competing AI labs.

Model distillation generally involves training a smaller model to reproduce knowledge or outputs from a larger model.

ByteDance founder Zhang Yiming believes independent development is necessary if the company wants to eventually outperform its competitors, according to the report. During an internal meeting two weeks ago, Zhang reportedly told the Seed team to focus on achieving world-leading AI capabilities over the long term rather than worrying about temporarily falling behind competitors.

ByteDance did not respond to a request for comment on the report, while Reuters said it had not independently confirmed the information.

The post ByteDance Working on Gigantic AI Model to Take on Anthropic’s Best appeared first on ProPakistani.

Show More

Related Articles

Leave a Reply

Your email address will not be published. Required fields are marked *

Back to top button

Adblock Detected

Please consider supporting us by disabling your ad blocker