AI Research Intern
Opusclip
| Company | Opusclip |
| Category | Data & Analytics |
| Location | Mountain View |
| Remote | On-site (inferred) |
| Employment | Internship |
| Level | Intern |
| Salary | Not stated by the employer |
| Posted | 23 Jul 2026 |
| Last verified | 12 Aug 2026 |
| Source | The employer's own careers page (company_site) |
Description
OpusClip is an AI video agent company that has raised $50 million and serves over 10 million creators. This AI Research Intern role involves exploring cutting-edge multimodal AI, LLMs, computer vision, speech, and agent systems to develop research prototypes and ship features across OpusClip and AgentOpus products.
What You'll Do
• Research and develop deep learning models in computer vision, speech & audio, multimodal understanding/generation, or LLM post-training
• Build AI-powered product features by integrating frontier foundation models into production systems using prompt engineering and Agent workflows
• Design scalable evaluation pipelines and domain-specific benchmarks for multimodal AI systems using automated evaluation methods
• Reproduce state-of-the-art research and translate advances into production-ready systems
• Collaborate with product, engineering, and AI research teams to rapidly prototype and ship new capabilities
What You Need
• Currently pursuing or recently completed Master's degree in Computer Science, Artificial Intelligence, Mathematics, or related field
• Solid understanding of Transformer architecture, Attention mechanisms, and familiarity with generative model families (GANs, diffusion models, autoregressive models)
• Strong Python programming skills with familiarity with Linux development environments, Git, and data structures
• Familiarity with media processing fundamentals including video and/or audio (e.g., ffmpeg, codecs, signal processing basics)
• Fluent in English with strong technical reading and writing skills, including ability to read research papers and write technical documentation
• Hands-on experience in one or more areas: low-level computer vision (Real-ESRGAN, SwinIR, BasicVSR++, diffusion-based SR), voice/speech (voice enhancement, voice cloning, voice generation), or LLM fine-tuning (SFT, RLHF, DPO, LoRA, PEFT)
Nice to Have
• Experience building Agent Systems or LLM-powered product features with frontier-model APIs (ChatGPT, Claude, Gemini)
• Familiarity with TypeScript
• Involvement in projects from inception to completion with strong coding fundamentals; open-source contributions
• Academic background or interest in video understanding/generation, multimodal systems, agents, and model evaluation/benchmarking
• Publications or involvement in academic research at top-tier venues (CVPR, ICCV, ECCV, NeurIPS, ICML, ICLR, ICASSP, Interspeech, AAAI, MM, TIP, TPAMI, ACL, EMNLP)