PodBrowser
Latent Space

A Technical History of Generative Media

Friday, 5 September 2025 · 6 min read · Listen to the episode ↗

The discussion features the evolution of generative media, with FAL's significant advancements in AI-driven image and video models, including notable releases like Stable Diffusion 1.5. There is a focus on community engagement and unique model development, addressing market gaps and optimizing performance through innovations such as custom kernels. The conversation highlights revenue implications from video models and anticipates AI's role in shaping future content creation, with expectations that a majority of video content could soon be AI-generated.

Alessio from Kernel Labs hosts a discussion with Gurkham and Batuhan from FAL, focusing on the evolution of generative media. Gurkham shares his transition from Amazon to FAL, where he developed a Python runtime that evolved into a generative media platform. Batuhan, head of engineering, highlights his role in building the Python cloud and their shared heritage, which fostered their collaboration. FAL has raised $125 million and attracted over 2 million developers, hosting around 350 models for image, video, and audio, generating over $100 million in revenue. Their strategy emphasizes unique models that address market gaps, relying on community feedback and internal evaluations.

The team actively engages with the community through platforms like Twitter and Reddit, monitoring trends around weekly model releases. Notable models include Stable Diffusion 1.5, which catalyzed FAL's entry into generative media, and STXL, which marked a revenue milestone. The introduction of Flux models and partnerships for video models further expanded their market presence, with the launch of VO3 enhancing user experience through Text-to-Video capabilities. FAL's decision to specialize in diffusion and inference was strategic, avoiding competition in language models against giants like OpenAI and Google.

The conversation reflects on the economic and creative implications of generative media, noting a significant shift in the AI landscape. The team has made substantial advancements in optimizing model performance, achieving significant reductions in processing time and enhancing efficiency through innovations like PyTorch 2.0 and custom kernels. The necessity for custom kernels tailored to various RMS norms is emphasized, resulting in over 100 custom kernels, enabling the generation of thousands of kernel shapes and significantly enhancing model performance.

The discussion contrasts image and language models, noting that image responses cannot be streamed like text. Latency is identified as vital for customer engagement, with evidence showing that slower latency adversely affects user metrics. The platform's success is partly attributed to the timing of open model releases for diffusion, which faced limited competition initially. Collaboration with closed-source model developers allows for high performance without code exposure. The inference engine is designed for self-service, enabling developers to deploy code and model weights without review.

The conversation addresses the challenges of serverless GPUs and the technology stack necessary for effective scaling, including a multi-cloud strategy involving six providers and 24 data centers. The majority of workloads are on H1N1 GPUs due to cost-effectiveness, while a dedicated team works on custom kernels for Blackwell GPUs. Ongoing debates about the efficiency of MMDITs versus single-stream models reflect diverse opinions in research teams. Distillation techniques, while initially promising, have not achieved long-term user retention, and the potential for a two-stage drafting process is explored.

The conversation emphasizes the need for effective image-to-image models, particularly in drawing, and discusses the role of editing models and control nodes like STXL. The decline in popularity of tools such as Fluxer raises questions about their ongoing relevance. One participant expresses a preference for cloud 4.1 for its quality, despite its slower speeds compared to Sonnet, highlighting the critical importance of fast draft generation for creators, especially in video models where speed has significantly improved.

The discussion also touches on autoregressive models and the early contributions of Gemini in the field, while OpenAI's 4.0 Image Generation is recognized as a major advancement. The rapid advancements in image generation are noted, with competitors quickly catching up after initial releases. There is excitement about new capabilities in video models and their potential applications in content creation, alongside speculation about mainstream adoption and affordability.

Skepticism is expressed regarding the effectiveness of world models in simulating reality, despite their ability to create consistent environments. Optimism exists in the robotics field, with expectations that advancements in data handling will enhance robotics models. The challenges of simulating gravitational forces are compared to the limitations of language models in basic arithmetic. Revenue growth from video models is highlighted, with one speaker noting an increase from 18% to over 50% attributed to advancements in editing models.

The impact of open-source models, particularly from Alibaba, is discussed, emphasizing their role in improving model quality and speed. The competitive landscape is noted, with Alibaba and smaller labs like Stepfun making significant strides in image and video model development. The conversation also addresses the attention garnered by training video models compared to language models, with a trend towards image and video models gaining prominence.

The dynamics of the market are explored, with a focus on revenue models and strategies, including various licensing approaches. The discussion on model usage reveals a power law in usage patterns, with users gravitating towards high-end video models or cost-effective alternatives. The conversation concludes with a focus on content moderation, noting that the generation of NSFW content is negligible, and the increasing enterprise usage of models for applications like chatbots and image generation tools.

Startups are investing heavily in launch videos, despite the availability of generative video technology that could reduce costs. The speaker believes we are in the early stages of generative media, with the potential for rapid advancements in the next 6 to 12 months, speculating that 80-90% of video content may soon be AI-generated and indistinguishable from traditional video. The speaker discusses the differences between ASLOP and product ads, noting that ASLOP creation is less concerned with appearance, while product ads require pixel-perfect accuracy.

Companies are investing in post-training for video models, with opportunities emerging for lip sync models and various video effects. Anticipation exists for new companies focused on post-training of open-source video models. Confi UI is highlighted as a community-driven platform with unique workflows, allowing model chaining through its File Workflows product. As models improve, workflows for images have become simpler, while video workflows remain complex.

There is encouragement for more model companies to raise funds and train models, emphasizing the need for data collection and prepared datasets for video models. Interest is expressed in building a startup focused on creative applications of AI in advertising, recognizing opportunities in targeted applications for specific industries. The unexpected popularity of image editing models and their rapid development is noted, alongside a market gap for affordable video models that excel in conversation and sound.

The discussion includes considerations on whether to integrate audio and video models or keep them separate, with a focus on timing and delivery. One speaker believes the current workflows and ConfiUI are nearly complete, with VO3 almost at full capacity. The company has rapidly expanded its workforce, actively seeking top talent in various roles. The engineering team is structured with a performance team overlapping with the applied ML team, which focuses on productionizing models and assisting customers.

The development of higher-level components aims to create easily integrable solutions, such as a virtual try-on feature for e-commerce. The conversation also explores the definition of "cracked" talent, particularly in technical problem-solving, with an example of a challenging task involving a sparse attention kernel. The recruitment strategy has successfully targeted applied ML engineers from platforms like Discord and Twitter, focusing on candidates passionate about generative media.

This summary was generated from the episode transcript and can contain mistakes.