Projects
Newest near the surface, older work deeper down.
5 mSafe Harbor
April 2026. DataHacks 2026, 48 hours, with Andres Rodriguez. I came up with the idea.
The question
With California sea temperatures rising, where will the animals go? The coast is defined by the creatures that live in it, and warming water is already pushing them around. If we can predict where temperature changes send a species, people can plan for irregular migration instead of reacting to it.
What it does
Safe Harbor is an interactive forecast of how rising sea surface temperatures could shift the habitat of two San Diego species, the garibaldi and the leopard shark. A slider moves the map from today out to two years and shows where each species is most likely to be found.
How we built it
We worked in a marimo notebook with Argo float data (18 ocean temperature profiles near San Diego) and about 3,000 iNaturalist sightings: 1,337 garibaldi and 1,665 leopard sharks. We mapped the temperatures and sightings together, then trained a random forest model on location, month, and temperature to predict where each species is likely to be. To look ahead, we estimated the warming trend from the float data, used bootstrapping to put a confidence interval on it, and shifted the temperatures forward in time.
- Python
- marimo
- scikit-learn
- NumPy
- pandas
- SciPy
- Matplotlib
- Cartopy
What it found
The float data showed warming of about 0.11 °C per year, with a 95% bootstrap interval from -0.008 to +0.238 °C per year. The interval just includes zero, so the trend is suggestive rather than conclusive. The current mean surface temperature in the area is 15.93 °C.
Limitations
Safe Harbor is a 48-hour prototype, and its results should be read that way.
- The warming trend comes from only 18 float profiles taken at different places and times, and its confidence interval includes zero.
- Sightings come from iNaturalist, so they reflect where people log animals as well as where the animals live. The "absence" points the model learns from are random spots in the ocean, not confirmed absences.
- The temperature changes we project over two years are small, so the differences between time horizons are small too.
- We did not test the model on held-out data, so we have no measure of how accurate its predictions are.
What was hard
The datasets were huge, so shrinking them and cleaning them down to San Diego took real time. We also had to learn Python tools we had never used, then build visualizations and a clean interface for the results.
What I learned, and what's next
Scope a project to the time you have, and pick up unfamiliar tools quickly. This was my first hackathon, and I was a freshman. Next steps are scaling the model and adding variables like water depth and more species.
110 mAI Image Detector
Summer 2025. SPIS at UCSD, 2 weeks, with Andres Rodriguez.
What it does
You give it any image and it answers AI or real, along with a confidence score. If the model is less than 70% confident, it says so instead of presenting the guess as a firm answer.
How it was built
A small convolutional neural network in PyTorch, trained in Google Colab on a Kaggle dataset of about 60,000 AI-generated and real images, all resized to 128 by 128. The network stacks four convolution blocks with batch normalization and dropout, then pools the result down to a final two-class prediction. I wrote it with a lot of AI-assisted coding.
- Python
- PyTorch
- torchvision
- Google Colab
How well it works
After 3 epochs it scored 69.87% on a held-out test set of 12,000 images. Validation accuracy was still rising at the end (56.6%, then 67.8%, then 69.7%), so more training is the obvious next step.
Limitations
This was the first project for both Andres and me, built in two weeks, and it has real weaknesses.
- It isn't reliable. It scored about 70% on the test set and closer to 60% in my own testing.
- It sometimes labeled images that were clearly AI-generated as real.
- It's a small network trained for only 3 epochs, on images shrunk to 128 by 128 pixels, which throws away fine detail.
- It was trained on a single Kaggle dataset, so it may not recognize images from other AI generators.
- Because so much of it was AI-assisted, I don't fully understand every part of how it works yet.

