Allocate is a private-markets platform that curates opportunities, creates feeder and white-label funds, automates KYC, subscriptions and capital calls, and provides AI-driven diligence, administration and portfolio tracking for advisors.
About Allocate Allocate is transforming private market investing by enabling RIAs and family offices to seamlessly discover, model, and manage their private market exposure. Our platform combines curated fund and co-investment opportunities with institutional-grade infrastructure. Through a single, data-rich digital experience, clients access top-tier opportunities across venture capital, private equity, private credit, and other private asset classes—backed by powerful tracking, analytics, and administration tools. Job Description Allocate is looking for a Senior Data Engineer to help build out the data infrastructure that powers our analytics, reporting, and data-driven product features. As a fintech startup on a mission to make investing in top-tier private markets more accessible, we have a wealth of financial and investment data to harness. Our data lead has established the foundational architecture and strategy, and we are now looking for a strong senior engineer to help extend, scale, and harden it. In this role you will partner closely with our data lead to model core financial entities, integrate internal and external sources, and build the pipelines and infrastructure that let our engineering and product teams make informed decisions and ship compelling features. This is a hybrid position based out of our Palo Alto, CA office, where you will work alongside our backend team (C#/.NET) and frontend team (Node/Vue.js) to integrate data pipelines into our platform. If you are a hands-on engineer who wants to do high-impact data work in a collaborative startup environment, we want to hear from you. Responsibilities - Build and Extend Data Architecture: Build on and extend Allocate's data lakehouse on AWS, combining data lake storage and warehouse technologies to store diverse financial datasets. Contribute to our knowledge graph that models key relationships (investors, funds, companies, etc.) and to the vector database integration that stores embeddings for semantic search and retrieval across our AI agents, models, and providers. - Develop Data Pipelines: Create robust ETL/ELT pipelines to ingest, clean, and transform data from various sources (internal application data and third-party APIs). Ensure both batch processing and real-time data streaming are handled to support up-to-date analytics and recommendations. Build pipelines with an eye on scalability (able to handle increasing data volume and complexity) and reliability (proper error handling and monitoring). - Enable AI/ML Capabilities: Work closely with our data science and engineering team to provision the data and infrastructure needed for machine learning models and AI features. This includes preparing training datasets, setting up feature stores, and orchestrating workflows that feed LLM-based agents with the context they need (e.g. retrieving relevant data via vector similarity search). You will also help implement systems to serve AI model outputs (such as recommendations) back into the product in real time. - Engineering Excellence and Collaboration: Partner with our data lead and the broader engineering team to deliver data and AI infrastructure. Raise the bar through thoughtful code review, testing, and adherence to best practices, and help engineers who consume data in their services do so effectively. Work in cross-functional squads to incorporate data-driven features into the product roadmap, and share your expertise with peers as the team grows. - Infrastructure and DevOps: Collaborate with our DevOps engineers to deploy and maintain data services. Containerize and orchestrate data tools (using Docker/Kubernetes on AWS EKS) for production use. Implement CI/CD pipelines for data workflows so that changes to data processing or models are tested and deployed automatically. Monitor the health and performance of our data platforms (setting up alerts, dashboards) and be ready to troubleshoot and resolve issues in production to ensure uptime of critical data and AI services. - Continuous Improvement: Stay up to date with the latest in data engineering and AI, from new AWS offerings to open-source ML tools. Evaluate and recommend new technologies, for example assessing whether a stream processing platform like Kafka/Kinesis or an orchestration tool like Airflow could improve pipeline reliability. Challenge conventions and innovate: we encourage rethinking how things are done as we push to build a world-class, intelligent platform. What You'll Need to Succeed - Strong Data Engineering Experience: 5+ years of hands-on experience in data engineering (or related fields), including designing and building large-scale data pipelines and storage solutions. You should have taken projects through the full lifecycle from design to production deployment. - Cloud Proficiency (AWS): Strong experience working with AWS cloud services for data. You should be comfortable with tools like S3, EC2, ECS, EKS, Athena, Redshift, Glue, and Step Functions. Experience setting up infrastructure-as-code (Terraform/CloudFormation) for these services is a plus. - Database and Data Modeling Skills: Proficiency in SQL and relational database design. Able to design efficient schemas and optimize queries/indexes for performance. Experience building or working with data warehouses or lakehouses (e.g. Snowflake, Databricks Delta Lake) is highly desired. Familiarity with graph databases (Neo4j, AWS Neptune, etc.) and knowledge graph schemas will help you hit the ground running. - Programming Expertise: Fluency in at least one major programming language used in data engineering. Python is commonly used for data pipelines, and pandas/PySpark experience is valuable. We also value experience with TypeScript/Node.js in data contexts, since our stack leans toward modern web technologies. The ideal candidate can work across languages, for example writing a data API in C# or Node.js to interface with our backend while also crafti