TikTokBellevue, WA
TikTok logo

Software Engineer Intern, Trust & Safety — Feed Safety Core Infrastructure

May 2026 Jul 2026

  • Autorescan: Designed, developed, and owned the recovery/rescan service of TikTok's new video-moderation core infrastructure: on average ~281K videos/day rescanned, saving ~3100 moderator-hours/day. During a downstream outage, autorescan served as the primary recovery path — recovered 18M+ videos and contained the incident, preventing P0 escalation.
  • Primary On-Call: Served as primary on-call (with senior SRE escalation) for core video moderation infrastructure during the legacy→new architecture migration — 4 distinct P0s across US, EU, and RoW in 7 days; authored RCAs for all and presented a detailed postmortem of a major 40-min global outage, with action items, to 50+ engineers including skip-level leadership.
  • Incident SOP: Independently resolved a recurring EU-region cascading failure: identified the critical under-provisioned model and established the mitigation SOP during volatile traffic migration, cutting MTTR from ~2h (P0 incident) to sub-20min (preemptive).
  • AI Tooling: Packaged autorescan operations into an MCP server with reusable agent skills adopted across the team for on-call automation; built automated failure reporting that surfaces cases landing in human review due to model confusion — a major tuning signal for next-gen moderation models.
  • Patrol Agent: Built and deployed a Grafana-driven on-call agent (ReAct loop) that monitors per-MQ rate limits and preemptively recommends tuning — the first iteration of the new Stability SWAT team's on-call master bot.
  • Data Pipeline: Built an hourly Hive-to-ClickHouse sync pipeline for fast failure analytics queries — optimized Spark SQL jobs for throughput; designed schema, partitioning, and TTL policies.
GogRPCProtobufThriftHiveSparkClickHouseGrafanaMCPKafka

University of Michigan – Welch LabAnn Arbor, MI
University of Michigan – Welch Lab logo

Research Assistant, ML / Bioinformatics

Sep 2025 May 2026

  • Representation Learning: Built transformer-based, cell-type-aware embeddings from protein language models (ESM-2, ProteinBERT) with learned attention pooling and an ontology-aware hierarchical loss over the cell-type DAG.
  • McCells & Scipher: Owner of two model tracks (one CZI-sponsored), trained on tens of millions of single-cell profiles for hierarchical cell-type annotation.
  • Designed a hierarchical loss function using probability marginalization to enforce Cell Ontology constraints in a DAG structure, preventing double-counting of shared ancestors.
  • Built an ML pipeline processing millions of cells from CellxGene Census using memory-efficient SOMA streaming for hierarchical classification across 100+ cell types.
PyTorchCellxGene CensusTileDB/SOMAProntoCell OntologyPython

BigHat BiosciencesSan Mateo, CA
BigHat Biosciences logo

Software Engineer Co-op — Full-Stack, LIMS & Research Infrastructure

Jan 2025 Aug 2025

  • ReArray: Built a full-stack reagent-tracking and automated liquid-handling module for a high-throughput Laboratory Information Management System (LIMS) — tracking any reagent in any well and auto-computing dilutions and array layouts across variable-sized, multi-plate workflows.
  • Audit-Trail UI: Built a React/TypeScript + Python audit-trail UI over DynamoDB surfacing deleted-record history for a large team of scientists, cutting audit lookups from hours to seconds.
  • Serverless Pipelines: Maintained Lambda/Kinesis serverless event pipelines and optimized CI/CD to sub-5-min dependency-pinned Docker releases, keeping experimental-event flow traceable across the LIMS.
  • Scientific Computing: Refactored scientific-computing modules for picomolar-level precision, eliminating rounding errors across 20K+ assay calculations/day.
  • ORM Standardization: Led migration of 80+ legacy Pydantic and SQLAlchemy models with inconsistent schemas to a unified ORM pattern, executing zero-downtime database migrations via Alembic to enable FDA-compliant data traceability.
ReactTypeScriptPythonFastAPIAWSAWS lambdaAWS KinesisDynamoDBDockerAlembicSQLAlchemyPydanticPostgreSQLcypress

UCSF – Andrej Šali LabSan Francisco, CA
UCSF – Andrej Šali Lab logo

Research Assistant

Mar 2023 Sep 2023

  • Built Markov state models from molecular dynamics simulations to capture conformational dynamics of intrinsically disordered proteins, identifying key metastable states and transitions.
  • Automated GPU job management (Bash) for large-scale simulation/data pipelines and integrated FRET data to improve model precision for biomolecular condensates.
PythonbashGROMACSOpenMMPyTorchNumPyPandasMarkov State ModelsGPU Computing

Lawrence Berkeley National LaboratoryBerkeley, CA

Research Assistant

Aug 2022 Jan 2023

  • Contributed to a retrosynthesis algorithm, generating polyketide synthase sequences for target molecules — adopted by multiple LBNL research teams.
  • Upgraded the ClusterCAD website's Django backend modules to integrate the latest retrosynthesis methods.
PythonDjangoPostgreSQLReact

GeopogoBerkeley, CA

Software Engineer Intern

May 2022 Aug 2022

  • Developed AR/3D features in Unity for real-time architectural visualization.
UnityC#MagicLeapAR

Institute of Computing Technology, Chinese Academy of SciencesBeijing, China

Student Researcher

Jun 2019 Aug 2019

  • Built a mass-spectrometry comparison system to identify best-matched peptides from MS2 spectra.
  • Implemented a pipeline mapping experimental data to an existing database, accelerating peptide identification.
PythonShellBioinformaticsMass Spectrometry