LOADING

MiniMax H3, Seedance 2.5, DeepSeek V4 are all here, and domestic big models are a bit busy today

What day is it today? MiniMax H3, Seedance 2.5, DeepSeek V4 Flash have all been released!

First, MiniMax released the first open source multimodal generation model H3, which directly put video generation, editing and multimodal understanding into the same model. Then, ByteDance introduced a new generation of video creation model Seedance 2.5, which lengthened the duration of a single video generation to 30 seconds; On the other hand, DeepSeek is also busy, announcing the public beta of the V4 Flash API, with a focus on upgrading Agent and code capabilities.

Within one day, three models were showcased together. Two aimed at video creation, one delved into the developer’s code terminal, and the domestic large model manufacturers seemed to have reached an agreement and collectively submitted their assignments on the last day of July.

MiniMax H3: Video models are also pursuing ‘omnipotence’
In the past, video models usually had clear division of labor. Some are responsible for creating videos, some are skilled in creating videos, and some specialize in video editing or action transfer. Users need to switch back and forth between multiple models and workflows to complete a piece.

H3 is attempting to dismantle these boundaries.

It supports different forms of input such as text, images, audio, and video, and understands users’ creative intentions in the same multimodal context before completing generation, editing, and expression. The model can directly produce content with a resolution of 2K, and can generate up to 15 seconds of audio and video at a time. It also focuses on improving command compliance, text and brand information presentation, as well as V2V Motion Transfer, which is the ability to transfer video to video actions.

For example, advertisers can have the model reference a real-life action while replacing characters, clothing, products, and scenes; E-commerce brands can also input product images, brand fonts, reference videos, and music to directly generate advertising materials that include product displays, character performances, and brand information.

From the ranking results, H3 currently ranks first in the Artificial Analysis video editing list with audio, and is also at the forefront of the Wensheng Video and Tusheng Video projects.

 

In terms of price, H3 is also quite competitive: the price for generating 2K videos is 0.8 yuan per second, which is about one-third of similar flagship models in the industry. In order to lower costs, MiniMax uses a high compression tokenizer to reduce the number of tokens required for video generation, and optimizes heterogeneous training, task load balancing, and GPU utilization efficiency.

At the same time, H3 will soon open up model weights, becoming MiniMax’s first open-source multimodal generative model. Enterprises can deploy locally and adapt to their own data and business scenarios; Domestic chip manufacturers also have the opportunity to optimize software and hardware around the model.

Seedance 2.5: Don’t just give me a piece of material, tell the story directly
In the past, using AI to generate videos often felt like drawing cards.

By inputting a prompt word, the model may generate a good shot, but changing the angle of the character is like changing their face. As soon as the scene switches, the clothing, lighting, and art style may also experience “amnesia”. To make a one minute video, creators often need to generate more than ten pieces of material, and then manually select, stitch, and process transitions.

The first thing Seedance 2.5 wants to solve is the problem of AI videos jumping out one by one.

This model continues the unified multimodal audio and video joint generation architecture of Seedance 2.0, and can generate up to 30 seconds of video at a time, which is twice as long as the previous generation’s 15 seconds. The upgrade here is not just about elongating one shot, but organizing multiple logically related scenes within 30 seconds, giving the story a foundation, progression, turning point, and ending.

If 30 seconds is not enough, users can continue generating videos after they already exist. The model will strive to maintain consistency in characters, scenes, style, sound, and narrative rhythm, ultimately forming a few minutes of content with relatively unified audio-visual language. The process of dismantling the lens, repeatedly pulling out the card, and manually splicing in the past has been further compressed.

Seedance 2.5 is still striving to get rid of the lingering “greasy feeling” of AI videos. The skin of the character is no longer like it has been waxed three times, with eye contact, texture, light and shadow, and image saturation closer to real shots, while reducing the situation where the model adds subtitles and background music on its own.

Another upgrade is’ reference capability ‘. Users can input up to 30 images, 10 videos, and 10 audios at a time, allowing the model to simultaneously reference characters, sounds, compositions, scenes, props, and camera movements from different materials. Simply put, in the past, we used to send AI a reference image, but now we can throw it character biographies, scene images, action samples, and background music together.

The editing ability of Seedance 2.5 has also been further enhanced. Users can specify the second of the plot, the action to be taken, and the perspective to be used through timestamps, or they can only modify characters, voices, or camera movements in specific segments. For example, retaining the actor’s movements and only replacing the green screen with another scene; Alternatively, the characters and images may remain unchanged, and only the camera movements from the 5th to the 10th second may be redesigned.

This means that Seedance is gradually moving from a ‘video material generator’ to a creative tool with directing, photography, and editing capabilities.

DeepSeek V4 Flash: The model remains unchanged, but it was’re educated ‘during post training
On July 31st, the official version of DeepSeeker V4 Flash API began public testing. Users can use the latest version by keeping the original API calling method and setting the model name to deepseek-v4-flash.

DeepSeeker V4-Flash-0731 has the same model architecture and size as the previous preview version, and the performance improvement mainly comes from retraining. In other words, DeepSeek did not give the model a bigger brain, but instead sent the same brain back to the training camp, focusing on tool calling, task planning, code development, and agent execution capabilities.

According to data released by DeepSeek, V4 Flash achieved significant improvements in Terminal Bench 2.1, NL2Repo, DeepSWE, Toolathlon, and multiple Agent tests, with Terminal Bench 2.1 scoring 82.7. It should be noted that DSBench FullStack and DSBench Hard belong to the internal testing set of DeepSeek, and the relevant results still need to be verified by more external results.

 

The official version V4 Flash supports the Responses API and has been adapted for Codex. Developers can call the DeepSeek model in the agent development environment such as Codex to let the model complete code reading, task disassembly, file modification, tool call and result check, rather than just answer an isolated programming question. The Harness framework used by DeepSeek to complete code agent tasks will also be released in the future.

To some extent, the upgrade of V4 Flash also re emphasizes the value of post training. As the base model becomes larger and the cost of pre training increases, how to use high-quality data, reinforcement learning, and task environments to “tune” existing models into more reliable agents is becoming another competitive mainline.

However, only the V4 Flash API was upgraded today, and DeepSeek’s app, web version, and V4 Pro have not yet been updated synchronously. The official version of V4 Pro will be released later. Therefore, it is still too early to describe today as “Flash surpassing Pro in all aspects”. A more accurate statement is that the ability gap between lightweight models and flagship models is being rapidly compressed by post training.

© 版权声明

相关文章