5.2 KiB
CLAUDE.md
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
Repository overview
This repository currently contains the machine-learning pipeline for ZeaVis Edu: a corn leaf disease classifier trained with EfficientNetV2B0 and exported for production use as TensorFlow SavedModel, TFLite, and TensorFlow.js formats.
The active project lives under Machine_Learning/. Most commands should be run from that directory unless noted otherwise.
Common commands
cd Machine_Learning
Set up a Python environment:
python -m venv venv
source venv/bin/activate
pip install -r requirements.txt
Run local dataset preprocessing after placing dataset_1.zip, dataset_2.zip, and dataset_3.zip beside preprocessing.py:
python preprocessing.py
Export a trained Keras model to SavedModel and TFLite after placing the Colab-trained model at best_model/best_model.keras:
python save_model.py
Convert the SavedModel export to TensorFlow.js via CLI:
export PROTOCOL_BUFFERS_PYTHON_IMPLEMENTATION=python
tensorflowjs_converter \
--input_format=tf_saved_model \
--output_format=tfjs_graph_model \
--signature_name=serving_default \
--saved_model_tags=serve \
model/saved_model \
model/tfjs_model
Open the training notebook locally if needed:
jupyter notebook notebook.ipynb
There is no project test suite, lint command, or build system configured in the current repository.
Fullstack app commands
The TypeScript application scaffold lives at the repository root and uses Bun workspaces with Moon tasks.
Install dependencies:
bun install
Run all development tasks through Moon:
bun run dev
Run type checks:
bun run typecheck
Run production builds:
bun run build
Run the API directly:
cd apps/api && bun run start
Run the web app directly:
cd apps/web && bun run dev
High-level architecture
Machine_Learning/preprocessing.pyprepares the training dataset locally. It extracts three source ZIP files, merges selected class folders intodataset/, maps selected Mandarin labels from Dataset 3 viadesc.json, removes known problematic image files, then createsdataset.zipfor upload to Google Drive/Colab.Machine_Learning/notebook.ipynbis the training workflow intended for Google Colab with GPU enabled. It trains an EfficientNetV2B0-based classifier and saves the best model to Google Drive asbest_model.keras.Machine_Learning/save_model.pyis the production export step. It loadsbest_model/best_model.keras, rebuilds a clean EfficientNetV2B0 architecture without training-time augmentation layers, copies weights into that model, exportsmodel/saved_model/, and writesmodel/model.tflite.- TensorFlow.js export is intentionally done with the
tensorflowjs_converterCLI rather than from Python to avoid protobuf/runtime conflicts documented in the README.
Fullstack application architecture
The root TypeScript workspace is a Bun + Moon monorepo:
apps/web/contains the React + Vite + TypeScript frontend with React Router, TanStack Query, Zustand, Tailwind, and shadcn/ui-style components.apps/api/contains the Elysia backend with health/status routes and Drizzle/PostgreSQL configuration.packages/shared/contains shared TypeScript types and utilities consumed by both apps.
The backend reads DATABASE_URL for Drizzle/PostgreSQL, but the initial health/status endpoints do not require a live database connection.
Model labels and dataset mapping
The classifier targets four Indonesian labels:
Bercak Daun— Gray Leaf SpotHawar Daun— Northern/Southern Leaf BlightKarat Daun— Common RustDaun Sehat— healthy corn leaf
Dataset handling is part of the model logic:
- Dataset 1 contributes
Bercak Daun,Hawar Daun, andDaun Sehat; itsKarat Daunfolder is intentionally ignored because the README states it is not representative. - Dataset 2 contributes
Common_Rustmapped toKarat DaunandHealthymapped toDaun Sehat. - Dataset 3 is routed through Mandarin label mappings in
PEMETAAN_KATEGORIinsidepreprocessing.py.
Important generated/local artifacts
The following files/directories are generated or externally supplied during the ML workflow and may not exist in a fresh clone:
Machine_Learning/dataset_1.zip,dataset_2.zip,dataset_3.zip— manually downloaded source datasets.Machine_Learning/dataset/andMachine_Learning/dataset.zip— generated bypreprocessing.py.Machine_Learning/best_model/best_model.keras— trained model downloaded from Colab/Google Drive.Machine_Learning/model/saved_model/,model/model.tflite, andmodel/tfjs_model/— production exports.
Notes for future changes
- Keep README command examples and this file in sync when changing the ML pipeline.
- Preserve the current class label names unless the training notebook, preprocessing mappings, and downstream app/API expectations are updated together.
save_model.pyassumes the clean architecture matches the trained model weights exactly; changes to the notebook model architecture usually require corresponding changes inbuild_clean_model().