08:15:13
TikTokBellevue, WA
Software Engineer Intern, Trust & Safety — Feed Safety Core Infrastructure
May 2026 – Jul 2026
- Autorescan: Designed, developed, and owned the recovery/rescan service of TikTok's new video-moderation core infrastructure: on average ~281K videos/day rescanned, saving ~3100 moderator-hours/day. During a downstream outage, autorescan served as the primary recovery path — recovered 18M+ videos and contained the incident, preventing P0 escalation.
- Primary On-Call: Served as primary on-call (with senior SRE escalation) for core video moderation infrastructure during the legacy→new architecture migration — 4 distinct P0s across US, EU, and RoW in 7 days; authored RCAs for all and presented a detailed postmortem of a major 40-min global outage, with action items, to 50+ engineers including skip-level leadership.
- Incident SOP: Independently resolved a recurring EU-region cascading failure: identified the critical under-provisioned model and established the mitigation SOP during volatile traffic migration, cutting MTTR from ~2h (P0 incident) to sub-20min (preemptive).
- AI Tooling: Packaged autorescan operations into an MCP server with reusable agent skills adopted across the team for on-call automation; built automated failure reporting that surfaces cases landing in human review due to model confusion — a major tuning signal for next-gen moderation models.
- Patrol Agent: Built and deployed a Grafana-driven on-call agent (ReAct loop) that monitors per-MQ rate limits and preemptively recommends tuning — the first iteration of the new Stability SWAT team's on-call master bot.
- Data Pipeline: Built an hourly Hive-to-ClickHouse sync pipeline for fast failure analytics queries — optimized Spark SQL jobs for throughput; designed schema, partitioning, and TTL policies.
GogRPCProtobufThriftHiveSparkClickHouseGrafanaMCPKafka
University of Michigan – Welch LabAnn Arbor, MI
Research Assistant, ML / Bioinformatics
Sep 2025 – May 2026
- Representation Learning: Built transformer-based, cell-type-aware embeddings from protein language models (ESM-2, ProteinBERT) with learned attention pooling and an ontology-aware hierarchical loss over the cell-type DAG.
- McCells & Scipher: Owner of two model tracks (one CZI-sponsored), trained on tens of millions of single-cell profiles for hierarchical cell-type annotation.
- Designed a hierarchical loss function using probability marginalization to enforce Cell Ontology constraints in a DAG structure, preventing double-counting of shared ancestors.
- Built an ML pipeline processing millions of cells from CellxGene Census using memory-efficient SOMA streaming for hierarchical classification across 100+ cell types.
PyTorchCellxGene CensusTileDB/SOMAProntoCell OntologyPython
BigHat BiosciencesSan Mateo, CA
Software Engineer Co-op — Full-Stack, LIMS & Research Infrastructure
Jan 2025 – Aug 2025
- ReArray: Built a full-stack reagent-tracking and automated liquid-handling module for a high-throughput Laboratory Information Management System (LIMS) — tracking any reagent in any well and auto-computing dilutions and array layouts across variable-sized, multi-plate workflows.
- Audit-Trail UI: Built a React/TypeScript + Python audit-trail UI over DynamoDB surfacing deleted-record history for a large team of scientists, cutting audit lookups from hours to seconds.
- Serverless Pipelines: Maintained Lambda/Kinesis serverless event pipelines and optimized CI/CD to sub-5-min dependency-pinned Docker releases, keeping experimental-event flow traceable across the LIMS.
- Scientific Computing: Refactored scientific-computing modules for picomolar-level precision, eliminating rounding errors across 20K+ assay calculations/day.
- ORM Standardization: Led migration of 80+ legacy Pydantic and SQLAlchemy models with inconsistent schemas to a unified ORM pattern, executing zero-downtime database migrations via Alembic to enable FDA-compliant data traceability.
ReactTypeScriptPythonFastAPIAWSAWS lambdaAWS KinesisDynamoDBDockerAlembicSQLAlchemyPydanticPostgreSQLcypress
UCSF – Andrej Šali LabSan Francisco, CA
Research Assistant
Mar 2023 – Sep 2023
- Built Markov state models from molecular dynamics simulations to capture conformational dynamics of intrinsically disordered proteins, identifying key metastable states and transitions.
- Automated GPU job management (Bash) for large-scale simulation/data pipelines and integrated FRET data to improve model precision for biomolecular condensates.
PythonbashGROMACSOpenMMPyTorchNumPyPandasMarkov State ModelsGPU Computing
Lawrence Berkeley National LaboratoryBerkeley, CA
Research Assistant
Aug 2022 – Jan 2023
- Contributed to a retrosynthesis algorithm, generating polyketide synthase sequences for target molecules — adopted by multiple LBNL research teams.
- Upgraded the ClusterCAD website's Django backend modules to integrate the latest retrosynthesis methods.
PythonDjangoPostgreSQLReact
GeopogoBerkeley, CA
Software Engineer Intern
May 2022 – Aug 2022
- Developed AR/3D features in Unity for real-time architectural visualization.
UnityC#MagicLeapAR
Institute of Computing Technology, Chinese Academy of SciencesBeijing, China
Student Researcher
Jun 2019 – Aug 2019
- Built a mass-spectrometry comparison system to identify best-matched peptides from MS2 spectra.
- Implemented a pipeline mapping experimental data to an existing database, accelerating peptide identification.
PythonShellBioinformaticsMass Spectrometry