New Deepseek, Seedance 2.5, Minimax H3, Gemini Robotics, AMD models: AI NEWS

This week in AI saw major advancements including Netflix’s IDV for video style editing, Deepseek’s efficient V4 Flash model, and cutting-edge video generators Seed Dance 2.5 and Miniax H3, alongside Google’s Gemini Robotics 2 enhancing robot dexterity. Additionally, AMD entered AI training with the Instellae model, while new tools like Redesign and Aiogram’s Object Remover improved image editing, highlighting significant progress across video, robotics, transcription, and multimodal AI technologies.

This week in AI has been packed with groundbreaking releases and updates across various domains. Netflix introduced IDV, an open-source AI that can change the style of a video scene without altering character identity or movement, enabling users to edit key frames and propagate style changes throughout the video. Additionally, Crisper Whisper 2, a powerful open-source transcription tool, was launched, offering verbatim and polished transcript modes with precise word-level timing, outperforming many existing transcription models. Deepseek released its latest V4 Flash model, which rivals top models like GLM and Opus in performance but is significantly smaller and cheaper, making high-level AI intelligence more accessible locally.

In the video generation space, Bite Dance unveiled Seed Dance 2.5, currently the best video model for high-action scenes and character consistency, supporting multimodal inputs and generating videos up to 30 seconds long at 720p, with plans for higher resolutions soon. Miniax followed with their H3 model, a flexible multimodal video generator capable of producing 2K resolution videos and supporting various input types, including text, images, video, and audio. Notably, Miniax H3 is set to be open-sourced soon, offering a more affordable alternative to Seed Dance 2.5. Both models represent significant advancements in AI-driven video creation, enabling more detailed and versatile content generation.

On the robotics front, Google DeepMind released Gemini Robotics 2, a comprehensive model that controls robots from walking to fine finger movements, allowing humanoid robots to perform complex tasks like picking up objects and manipulating them with dexterity. This update includes multiple models, including an offline version for local robot deployment, enhancing real-world robotic applications. Complementing this, a new system called Prism was introduced to improve robot control by integrating multiple sensor inputs for more informed and successful physical interactions, with code available for public use.

AMD made a notable entry into AI model training with Instellae, an open-source mixture of experts model trained entirely on AMD hardware, breaking Nvidia’s dominance in the space. This 16 billion parameter model is efficient and reportedly outperforms similar-sized models, with training checkpoints and code openly available. Meanwhile, Thinking Machines released Inkling Small, a multimodal open-source model capable of understanding text, audio, images, and video, offering strong performance with fewer active parameters, making it a cost-effective option for multimodal AI tasks.

Other exciting developments include AI tools like Redesign, which converts flat images into editable layers for easier customization, and Aiogram’s Object Remover, a free online tool for removing unwanted objects from images with high accuracy. Additionally, new interactive video world models like Wonder and Fi0 enable real-time exploration and physical reasoning in generated video scenes. Google also enhanced its Gemini app with AI-powered voice transcription and editing for Mac users. Overall, this week showcased remarkable progress in AI models, tools, and applications across video, robotics, transcription, and image editing, promising more accessible and powerful AI capabilities ahead.