Interhuman AI builds social intelligence infrastructure that makes AI systems capable of understanding human behavior. Our API detects and interprets behavioral signals like hesitation, engagement, confusion, interest - across voice, facial expressions, body language, and words. Unlike emotion AI, we focus on observable behaviors and their meaning in context.
We're looking for a motivated intern to join our AI and data team in Copenhagen. This is a hands-on role where you'll work directly with our AI engineers on the data infrastructure that powers our social intelligence API.
What You’ll Do
- Dataset creation and management: Support the creation of new training datasets by identifying relevant sources, structuring data pipelines, and documenting dataset specifications. You'll work closely with our ML team to understand what makes a high-quality dataset.
- Pipeline optimization: Help refine our dataset filtering pipeline by testing different parameters, identifying edge cases, and suggesting improvements. You'll gain hands-on experience with data processing tools and learn how to optimize for both quality and scale.
- Source investigation: Research and evaluate new potential data sources for behavioral signals. This includes investigating public datasets, APIs, and other resources that could enhance our training data coverage across different contexts and demographics.
- Quality assurance: Conduct quality checks on processed datasets, flag anomalies or inconsistencies, and help establish best practices for data validation. Your attention to detail will directly impact model performance.
Who we’re looking for
You might be a great fit for this role if you meet some or all of the following:
- Technical proficiency: You're comfortable with scripting (Python, JavaScript, or similar) and can quickly prototype solutions to automate data processing tasks. You're handy with code and can turn ideas into working scripts efficiently.
- Educational background: You're currently pursuing or have completed a degree in Computer Science, Data Science, Machine Learning, or a related technical field.
- Dataset curation expertise: You've worked on curating datasets before and understand the importance of data quality, consistency, and proper documentation in building robust AI systems.
- Experience with constrained optimization: You have practical experience working with optimization problems under constraints, whether in machine learning, data processing, or related fields.
- Vision-Language Model (VLM) fine-tuning experience: Experience fine-tuning VLMs would be a strong plus, as it directly relates to our multimodal approach to understanding behavioral signals.
- Self-directed and resourceful: You're comfortable figuring things out independently and don't need constant direction. When you encounter obstacles, you find creative solutions and know when to ask for help.
- Strong communication skills: You can explain technical concepts clearly in writing and speech.
- Thrive in a fast paced environment: You are good at prioritizing and balancing multiple tasks, and flexible when changes occur.