Welcome to your interview.
The question is on your right: Build a Multimodal LLM for Real-Time Video Editing. Take a moment with it first.
Talk your thinking through with me if you like - when you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes). Discussion and graded submissions share your five interviewer interactions, so spend them well.
VideoPro, a leading video editing software provider with over 1 million active users, aims to integrate AI-driven features that enhance user creativity and efficiency. The goal is to build a multimodal large language model (LLM) that can understand user commands in natural language and interact with video content in real-time, enabling seamless editing workflows.
| Feature Group | Count | Examples |
|---|---|---|
| Video Metadata | 10 | duration, resolution, frame_rate, codec |
| User Commands | 5 | 'Trim video', 'Add filter', 'Speed up', 'Add text overlay' |
| Video Content | 50K | frames, audio segments, color histograms, scene descriptors |
| User Interaction | 15 | click events, time spent on each tool, undo actions |