MiniCPM-o 4.5: A 9B Model That Can See, Hear, and Speak at the Same Time
Most multimodal AI models today work in turns. You feed them a video, wait for them to process it, and then get a response. While this works well for traditional chatbots, it does not reflect how humans naturally interact with the world. We listen while we speak. We react to things as they are still…