Dhinesh Kumar Ravi
Poster Presenter
Dr. Dhinesh Kumar R is a Principal AI Architect at OO Studio AI, India, with extensive experience in Artificial Intelligence, Edge Computing, Intelligent Transportation Systems, and Next-Generation Vehicular Communication (V2X). He holds a PhD in Edge AI-based context-aware Quality of Service (QoS) optimization for V2X communications from Vellore Institute of Technology (VIT), India, and a Master’s degree in Mobile and Satellite Communication from the University of Surrey, UK. His research expertise spans Edge AI, Multi-Agent Reinforcement Learning, LLMs, Intelligent Transportation Systems (ITS), UAV networks, and 5G/6G communication systems, with a strong focus on real-time deployment in smart mobility and smart city environments. Dr. Dhinesh has contributed extensively to both academia and industry through roles at VIT’s Autonomous Vehicle Research Lab, IIT Roorkee (iHUB DivyaSampark), University of Petroleum and Energy Studies, and industry organizations including Mirabilis Design Inc. and DSTL (UK). He has led several cutting-edge projects in adaptive traffic control systems, UAV-based networks, and AI-driven V2X communication frameworks. He has published multiple high-impact journal papers in reputed journals such as Vehicular Communications, Results in Engineering, and Sustainable Energy Technologies and Assessments, and has presented research at leading IEEE and ACM international conferences. He also holds multiple published patents in adaptive vehicular communication systems, UAV optimization, and intelligent traffic control systems. His technical expertise includes Python, MATLAB, C, TensorFlow, PyTorch, AWS, Azure, NS-3, OMNeT++, and embedded platforms such as Jetson Nano, Jetson Xavier NX, Raspberry Pi, and Coral TPU. He actively works across domains including Edge/Cloud AI systems, autonomous vehicles, wireless networks, and agent-based learning systems. Dr. Dhinesh has also served as a technical reviewer for leading journals including IEEE Transactions on Fuzzy Systems, Alexandria Engineering Journal, and Science China Information Sciences. He is deeply involved in research collaboration, innovation-driven development, and mentoring multidisciplinary teams, with a strong focus on building scalable AI systems for real-world deployment in transportation, smart cities, and autonomous systems.
From Tokens to Emotions: The Next Frontier of Foundation Models, Agentic Intelligence, and Generative Media Large Language Models have moved in five years from next-token predictors to multi-modal reasoners, and the trajectory now bends sharply toward agentic intelligence systems that plan, delegate, and act across tools, modalities, and humans. Yet the next true frontier is no longer context windows or larger parameter counts; it is expressive fidelity, the capacity of generative AI to preserve human emotion, intent, and cultural nuance as content crosses languages, modalities, and creative pipelines. This keynote charts that progression and presents the architecture, training methodology, and production deployment of LEMOO, the world's first Large Emotional Model for media, and Visonus, the agent-orchestrated super-app built on top of it. The session opens with a technical map of the current generative AI landscape - the shift from autoregressive transformers to flow-matching and diffusion-based generators, the maturation of RLHF and direct preference optimization, the rise of retrieval-augmented and tool-using agents, and the emerging discipline of multi-agent orchestration where specialized models coordinate under a planner. LEMOO is positioned within this arc: a transformer-based foundation model trained with dual-reference flow matching to disentangle speaker identity, prosody, and emotional state, refined through preference-based fine-tuning aligned to director and listener feedback. Word-level spline editing and phoneme adaptation transform the model into a directable instrument closing the gap between one-click TTS and pro-grade voice performance. The talk then unpacks Visonus as a deployed multi-agent system in which specialized AI agents — SonusDUB for multilingual voice cloning, VisoTXT for cultural-context translation through multilingual graph mappings, VisoSYNC for 4K and ProRes lip-sync, VisoSEAL for neural watermarking, and CelebVOX for voice-IP licensing coordinate across a single secure namespace to execute end-to-end media pipelines. Edge cloud co-inference, INT8 quantization, and streaming decoders compress latency to studio-grade tolerances while AES-256 encryption and traceable watermarks satisfy SOC-2 and GDPR. The developed architecture is grounded in a growing portfolio of deployed cinema. Representative engagements include the five-language release of Kanappa (2025) with a 60% reduction in dubbing cost and time, multilingual delivery of Jackie Chan's The Legend preserving Chinese-accented rhythm across Indian languages, and dialect-faithful synthesis for the Telugu film Meiazhagan alongside ongoing collaborations with major studios and production houses across the Indian and global entertainment landscape. The session closes with the speaker's perspective on where the field is heading: emotionally aware foundation models as the successor to text-only LLMs, agent orchestration as the operating system of the next AI decade, on-device personalization for sovereign creative identity, and verifiable provenance as the trust layer of the synthetic media economy.