
Job Description
About Juxta
Juxta builds real-world datasets used to train and improve AI systems. High-quality models start with high-quality data, and our data collection operations are a critical part of that process.
We’re looking for a Data Processing Intern to work hands-on with our data collection team, helping ensure that the data we collect in the field is accurate, consistent, and ready to be used for model training.
What You’ll Do
This is a hands-on role that combines field data collection, quality assurance, and data preprocessing. You’ll work directly with Juxta’s data collectors and accompany them to different collection locations.
Your responsibilities will include:
- Support field data collection: Travel with Juxta’s data collectors to different locations and help oversee collection sessions.
- Ensure collection quality: Verify that collectors are following the correct procedures, equipment is configured properly, and collected data meets Juxta’s quality standards.
- Identify issues in real time: Catch problems such as incorrect setups, missing data, inconsistent collection procedures, or corrupted files before they affect an entire collection session.
- Validate collected data: Review datasets after collection to confirm that files are complete, correctly organized, and usable.
- Preprocess data: Prepare raw collected data for downstream training workflows, including cleaning, organizing, formatting, filtering, and validating data.
- Document collection sessions: Maintain clear records of collection conditions, issues encountered, and any changes that may affect the resulting dataset.
- Improve our processes: Work with the Juxta team to identify recurring data-quality issues and improve collection and preprocessing workflows.
What We’re Looking For
- Strong attention to detail and an ability to spot inconsistencies
- Comfortable working with files, datasets, and basic technical workflows
- Ability to follow detailed data-collection protocols and ensure others follow them consistently
- Strong organizational and problem-solving skills
- Comfortable working both independently and directly with a field team
- Willingness to travel locally to different data collection locations
- Interest in AI, machine learning, data engineering, robotics, computer vision, or related fields
- Currently pursuing or recently completed a degree in Computer Science, Data Science, Engineering, Statistics, or a related technical field
Nice to Have
- Experience working with Python
- Familiarity with data cleaning or preprocessing using tools such as Python, NumPy, Pandas, or similar
- Experience working with large datasets, sensor data, images, video, audio, or other multimodal data
- Familiarity with machine learning training pipelines or dataset preparation
- Previous experience with field research, data collection, QA, or technical operations
What You’ll Learn
You’ll get direct exposure to the process of building real-world datasets for AI—from collecting raw data in the field to preparing it for model training. You’ll work closely with Juxta’s technical and operations teams and develop practical experience with data quality, preprocessing, AI training pipelines, and large-scale data operations.
This role is a strong fit for someone who enjoys being hands-on, cares deeply about getting the details right, and wants to understand how high-quality training data is actually created.
Optimize Your Resume for This Job
Get a match score and see exactly which keywords you're missing
Job Details
- Category
- Software
- Employment Type
- Internship
- Location
- San Francisco, CA
- Posted
- Compensation
- $2,000 - $4,000 per month
About Juxta
Building a GPS alternative 100x more powerful with no hardware needed.
Similar Software Roles



Found this role interesting?