ENFR
8news

Tech • IA • Crypto

TodayTopicsVideosCryptoArchivesFavorites

Build creative apps with the GenMedia suite (I/O Connect ’26)

9.4/10
GoogleGoogle for DevelopersJuly 6, 2026 at 06:14 PM27:25
Audio player
0:00 / 0:00

TL;DR

Google DeepMind showcased a suite of generative media tools that can transform full books into cohesive multimedia experiences combining images, video, music, and voice.

KEY POINTS

Shift toward generative media

Generative AI is moving beyond text into images, video, audio, and voice, reflecting how people naturally communicate through media. This shift positions “gen media” as the next major phase after large language models, with multimodal systems becoming central to everyday workflows.

Multimodal foundation with Gemini

Gemini was designed from the outset as a multimodal model capable of processing text, images, and audio within a single system. Its large context window allows entire documents, including full books, to be ingested and used as a basis for generating diverse media outputs.

End-to-end content generation from books

A full public-domain book such as The Wind in the Willows can be uploaded once and reused across tasks. The system can extract characters, summarize chapters, and generate structured prompts that feed downstream models, enabling automated illustration and storytelling pipelines.

Character and scene visualization

Image generation models like Nano Banana can create consistent character portraits and scene illustrations. By chaining requests through a stateful API, the system maintains stylistic coherence across multiple outputs, ensuring characters retain recognizable features throughout a story.

Stateful Interactions API

The newly introduced Interactions API allows developers to maintain context across multiple requests without repeatedly sending data. This reduces token usage and improves efficiency, especially for complex workflows involving many assets like images and prompts.

Video generation with contextual prompts

The Veo video model can animate still images into short clips. Results improve when prompts describe what happens after the depicted scene, rather than repeating the original image prompt, highlighting the importance of temporal context in video generation.

AI-generated music soundtracks

The Lyria music model produces instrumental or full-length tracks tailored to scenes or chapters. It supports control over structure, tempo, and style, enabling background scores that align with narrative tone, such as “cyberpunk” or “acoustic” themes.

Text-to-speech with multiple characters

A Gemini TTS model can generate narrated dialogue using only two base voices while simulating multiple characters through prompting techniques. Variations in tone and style create the illusion of a full cast, useful for audiobooks or dramatized storytelling.

Integrated multimedia output

By combining images, narration, and music, the system can assemble cohesive multimedia content. Even complex scripts for synchronizing these elements can be generated automatically, demonstrating how AI can orchestrate full production pipelines.

Emerging video editing model Omni

The upcoming Omni model introduces fine-grained video editing via prompts. Instead of regenerating entire clips, users can modify specific elements—such as voice, style, or inserted objects—while preserving the rest of the footage, enabling iterative “prompt-based editing.”

Scalability considerations

While chaining large contexts works in demonstrations, production systems require optimization, such as selectively loading only relevant assets. This highlights ongoing trade-offs between capability, cost, and performance in large-scale deployments.

CONCLUSION

Advances in multimodal AI are enabling fully automated media creation pipelines, signaling a shift from text-centric tools to integrated systems that generate complete visual, audio, and interactive experiences.

Full transcript

More from Google