4.2 KiB
CLAUDE.md
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
Repository overview
This repository currently contains the machine-learning pipeline for ZeaVis Edu: a corn leaf disease classifier trained with EfficientNetV2B0 and exported for production use as TensorFlow SavedModel, TFLite, and TensorFlow.js formats.
The active project lives under Machine_Learning/. Most commands should be run from that directory unless noted otherwise.
Common commands
cd Machine_Learning
Set up a Python environment:
python -m venv venv
source venv/bin/activate
pip install -r requirements.txt
Run local dataset preprocessing after placing dataset_1.zip, dataset_2.zip, and dataset_3.zip beside preprocessing.py:
python preprocessing.py
Export a trained Keras model to SavedModel and TFLite after placing the Colab-trained model at best_model/best_model.keras:
python save_model.py
Convert the SavedModel export to TensorFlow.js via CLI:
export PROTOCOL_BUFFERS_PYTHON_IMPLEMENTATION=python
tensorflowjs_converter \
--input_format=tf_saved_model \
--output_format=tfjs_graph_model \
--signature_name=serving_default \
--saved_model_tags=serve \
model/saved_model \
model/tfjs_model
Open the training notebook locally if needed:
jupyter notebook notebook.ipynb
There is no project test suite, lint command, or build system configured in the current repository.
High-level architecture
Machine_Learning/preprocessing.pyprepares the training dataset locally. It extracts three source ZIP files, merges selected class folders intodataset/, maps selected Mandarin labels from Dataset 3 viadesc.json, removes known problematic image files, then createsdataset.zipfor upload to Google Drive/Colab.Machine_Learning/notebook.ipynbis the training workflow intended for Google Colab with GPU enabled. It trains an EfficientNetV2B0-based classifier and saves the best model to Google Drive asbest_model.keras.Machine_Learning/save_model.pyis the production export step. It loadsbest_model/best_model.keras, rebuilds a clean EfficientNetV2B0 architecture without training-time augmentation layers, copies weights into that model, exportsmodel/saved_model/, and writesmodel/model.tflite.- TensorFlow.js export is intentionally done with the
tensorflowjs_converterCLI rather than from Python to avoid protobuf/runtime conflicts documented in the README.
Model labels and dataset mapping
The classifier targets four Indonesian labels:
Bercak Daun— Gray Leaf SpotHawar Daun— Northern/Southern Leaf BlightKarat Daun— Common RustDaun Sehat— healthy corn leaf
Dataset handling is part of the model logic:
- Dataset 1 contributes
Bercak Daun,Hawar Daun, andDaun Sehat; itsKarat Daunfolder is intentionally ignored because the README states it is not representative. - Dataset 2 contributes
Common_Rustmapped toKarat DaunandHealthymapped toDaun Sehat. - Dataset 3 is routed through Mandarin label mappings in
PEMETAAN_KATEGORIinsidepreprocessing.py.
Important generated/local artifacts
The following files/directories are generated or externally supplied during the ML workflow and may not exist in a fresh clone:
Machine_Learning/dataset_1.zip,dataset_2.zip,dataset_3.zip— manually downloaded source datasets.Machine_Learning/dataset/andMachine_Learning/dataset.zip— generated bypreprocessing.py.Machine_Learning/best_model/best_model.keras— trained model downloaded from Colab/Google Drive.Machine_Learning/model/saved_model/,model/model.tflite, andmodel/tfjs_model/— production exports.
Notes for future changes
- Keep README command examples and this file in sync when changing the ML pipeline.
- Preserve the current class label names unless the training notebook, preprocessing mappings, and downstream app/API expectations are updated together.
save_model.pyassumes the clean architecture matches the trained model weights exactly; changes to the notebook model architecture usually require corresponding changes inbuild_clean_model().