GPT on a Leash: Evaluating LLM-based Apps & Mitigating Their Risks
The task of testing and evaluating AI systems is extremely challenging, especially when it involves text and unstructured data. In the case of LLM-based applications, these challenges are magnified by the fact that there isn't "one correct answer" and by a combination of various external constraints such as topics that shouldn't be discussed. Speaker: Philip is the co-founder and CEO of Deepchecks. Philip is an experienced Data Scientist and in the past, he led a top-tier ML research gr
Philip, co-founder and CEO of Deepchecks, on the challenges of testing and evaluating LLM-based applications, where there is no single correct answer, and how to mitigate their risks under external constraints.
More from the studio
1:01:41SF DEMO NIGHT 🚀 (w/ The AI Collective)
Six startup teams present live AI demos spanning interactive avatars, executive communication, home services, autonomous-agent security, enterprise operations, and personal intelligence, followed by audience Q&A and community voting.
Dmytro Spodarets·Sep 18, 2026
17:58Interview with Keerti Melkote: Anyscale Azure Integration and the Future of Ray
At Ray Summit 2025, Dmytro Spodarets speaks with Keerti Melkote about Anyscale, Ray, and the new Azure integration. The interview covers enterprise AI workloads, GPU challenges, agentic systems, and the future of scalable AI infrastructure.
Dmytro Spodarets·Nov 24, 2025
38:39Manus vs OpenAI: How a Startup Is Beating the Giants in the AI Agent Race
Tao Zhang of Manus AI shares how the startup is taking on OpenAI in the agent race, revealing product insights, strategy, and the future of intelligent tools.
Dmytro Spodarets·Jul 21, 2025