Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -6,3 +6,6 @@ experiments/*/

# JetBrains IDE idea directories
.idea/

# Documentation
docs/build/
40 changes: 1 addition & 39 deletions README.md
Original file line number Diff line number Diff line change
@@ -1,44 +1,6 @@
# Synergistic Software Platform for AI, Physics Simulations, and Experiments (Synapse)

Synapse is a modular framework to build components of digital twins.
It supports operators of machines and experiments with ML-assisted predictions, trained on a combination of continously measured and simulated data.
Synapse embraces emerging [integrated research infrastructures](https://www.nersc.gov/what-we-do/computing-for-science/integrated-research-infrastructure), deploying a user-facing cloud service, using HPC/cloud compute, and exchanging modular components via container registries.

At the moment, Synapse uses the NERSC Spin (control, dashboard), NERSC Superfacility API (simulation submission & ML training on Perlmutter), and NERSC container registry.
Synapse is actively developed and broadened to an AI-accelerated, portable framework.

## Overview

Synapse is a software platform that enables physicists to couple experimental data, simulations, and machine learning (ML) models trained on experimental and simulation data.
As an example, the schematic below illustrates how Synapse is used at the Berkeley Lab Laser Accelerator Center (BELLA):

![Synapse overview](synapse_overview.png)

One of the main software components is the graphical user interface (GUI), which is deployed via [Spin](https://docs.nersc.gov/services/spin/) at NERSC.
The application requires access to various data and information sources, as described below.

### Displaying ML predictions

To display ML predictions, the application requires the following:

- **Experiment configuration file**: A YAML file named `config.yaml` stored in the root directory of an experiment's repository that defines the input, output, and calibration variables.
- **Simulation and experimental data points**: Each data point consists of values for the scalar inputs and outputs defined in the experiment configuration file.
Data points are stored in a [MongoDB](https://www.mongodb.com/) database, where each experiment is represented by a separate collection.
Experimental and simulation data points are stored in the same collection and are distinguished by the `experimental_flag` attribute.
- **ML models**: Machine learning models that interpolate between data points and are stored in [MLflow](https://mlflow.org/).
- **Simulation movies** (optional): For certain experiments, users can click on simulation data points to visualize simulation movies.
The corresponding MP4 files are stored in the Perlmutter shared file system at `/global/cfs/cdirs/m558/superfacility/simulation_data`.
This directory is mounted on the container image running on Spin.

### Launching ML training at NERSC

ML models can be trained by launching jobs on Perlmutter from the GUI, through the [NERSC Superfacility API](https://docs.nersc.gov/services/sfapi/). The application requires the following:

- **Superfacility API credential file**: Instructions on generating and uploading the credential file from the GUI are in [dashboard/README.md](dashboard/README.md).
- **Submission script**: The batch script [ml/training_pm.sbatch](ml/training_pm.sbatch) is copied into the container image pushed to the NERSC registry and deployed via Spin (see [dashboard.Dockerfile](dashboard.Dockerfile)). It serves as a template for Superfacility API job submission when users launch model training from the GUI.
- **Python scripts and configuration files**: These include [ml/train_model.py](ml/train_model.py), [ml/Neural_Net_Classes.py](ml/Neural_Net_Classes.py), and the experiment configuration file `config.yaml`.
They are copied into the container image pushed to the NERSC registry and deployed via Spin (see [dashboard.Dockerfile](dashboard.Dockerfile)).
When users launch model training from the GUI, these files are copied to the Perlmutter shared file system at `/global/cfs/cdirs/m558/superfacility/model_training/src/` for access by the Superfacility API job. The `config.yaml` file is automatically populated with the configuration values specified in the GUI before being copied to the shared file system.
Synapse is documented in `docs/`. Start with the [overview](docs/source/overview.md), then see [dashboard](docs/source/dashboard.md), [ML training](docs/source/ml-training.md), and [experiment configuration](docs/source/experiment-configuration.md).

## Copyright Notice and License Agreement

Expand Down
200 changes: 2 additions & 198 deletions dashboard/README.md
Original file line number Diff line number Diff line change
@@ -1,199 +1,3 @@
# Table of Contents
* [Overview](#Overview)
* [Run the Dashboard Locally](#Run-the-Dashboard-Locally)
* [Without Docker](#Without-Docker)
* [With Docker](#With-Docker)
* [Run the Dashboard at NERSC](#Run-the-Dashboard-at-NERSC)
* [Get the Superfacility API Credentials](#Get-the-Superfacility-API-Credentials)
* [For Maintainers](#For-Maintainers)
* [Generate the conda environment lock file](#Generate-the-conda-environment-lock-file)
* [Build and push the Docker container to NERSC](#Build-and-push-the-Docker-container-to-NERSC)
* [References](#References)
# Dashboard

# Overview

The Synapse dashboard provides a web interface for working with data from experiments, simulations, and ML models.

The dashboard can be run in two distinct ways:

1. Locally on your computer.

2. At NERSC through Spin.

# Run the Dashboard Locally

This section describes how to develop and use the dashboard locally.

## Without Docker

### Prepare the conda environment

1. Move to the [dashboard/](./) directory.

2. Activate the conda environment `base`:
```bash
conda activate base
```

3. Install `conda-lock` if not installed yet:
```bash
conda install -c conda-forge conda-lock
```

4. Create the conda environment `synapse-gui`:
```bash
conda-lock install --name synapse-gui environment-lock.yml
```

### Run the dashboard

1. Create an SSH tunnel to access the MongoDB database at NERSC (in a separate terminal):
```bash
ssh -L 27017:mongodb05.nersc.gov:27017 <username>@dtn03.nersc.gov -N
```

2. Move to the [dashboard/](./) directory.

3. Set up the database settings (read-only) and the AmSC MLflow API key:
```bash
export SF_DB_HOST='127.0.0.1'
export SF_DB_READONLY_PASSWORD='your_password_here' # Use SINGLE quotes around the password!
export AM_SC_API_KEY='your_amsc_api_key_here' # Required when MLflow tracking_uri is AmSC
```

4. Activate the conda environment `synapse-gui`:
```bash
conda activate synapse-gui
```

5. Run the dashboard as a web application:
```bash
python -u app.py --port 8080
```

## With Docker

### Run the dashboard

1. Create an SSH tunnel to access the MongoDB database at NERSC (in a separate terminal):
```bash
ssh -L 27017:mongodb05.nersc.gov:27017 <username>@dtn03.nersc.gov -N
```

2. Move to the root directory of the repository.

3. Build the Docker image as described [below](#build-the-docker-image).

4. Run the Docker container:
```bash
docker run --network=host -v /etc/localtime:/etc/localtime -v $PWD/ml:/app/ml -e SF_DB_HOST='127.0.0.1' -e SF_DB_READONLY_PASSWORD='your_password_here' -e AM_SC_API_KEY='your_amsc_api_key_here' synapse-gui
```
For debugging, you can enter the container without starting the app:
```bash
docker run --network=host -v /etc/localtime:/etc/localtime -v $PWD/ml:/app/ml -e SF_DB_HOST='127.0.0.1' -e SF_DB_READONLY_PASSWORD='your_password_here' -e AM_SC_API_KEY='your_amsc_api_key_here' -it synapse-gui bash
```
Note that `-v /etc/localtime:/etc/localtime` is necessary to synchronize the time zone in the container with the host machine.

# Run the Dashboard at NERSC

Connect to the [dashboard](https://bellasuperfacility.lbl.gov/) deployed at NERSC through Spin and play around!
Remember that you need to upload valid Superfacility API credentials in order to launch simulations or train ML models directly from the dashboard.

# Get the Superfacility API Credentials

Following the instructions at [docs.nersc.gov/services/sfapi/authentication/#client](https://docs.nersc.gov/services/sfapi/authentication/#client):

1. Log in to your profile page at [iris.nersc.gov/profile](https://iris.nersc.gov/profile).

2. Click the icon with your username in the upper right of the profile page.

3. Scroll down to the section "Superfacility API Clients" and click "New Client".

4. Enter a client name (e.g., "Synapse"), choose `sf558` for the user, choose "Red" security level, and select either "Your IP" or "Spin" from the "IP Presets" menu, depending on whether the key will be used from a local computer or from Spin.

5. Download the private key file (in pem format) and save it as `priv_key.pem` in the root directory of the dashboard.
Each time the dashboard is launched, it will automatically find the existing key file and load the corresponding credentials.

6. Copy your client ID and add it on the first line of your private key file as described in the instructions at [nersc.github.io/sfapi_client/quickstart/#storing-keys-in-files](https://nersc.github.io/sfapi_client/quickstart/#storing-keys-in-files):
```
randmstrgz
-----BEGIN RSA PRIVATE KEY-----
...
-----END RSA PRIVATE KEY-----
```

7. Run `chmod 600 priv_key.pem` to change the permissions of your private key file to read/write only.

# For Maintainers

## Generate the conda environment lock file

1. Move to the directory [dashboard/](.).

2. Activate the conda environment `base`:
```bash
conda activate base
```

3. Install `conda-lock` if not installed yet:
```bash
conda install -c conda-forge conda-lock
```

4. Generate the conda environment lock file:
```bash
conda-lock --file environment.yml --lockfile environment-lock.yml
```

## Build and push the Docker container to NERSC

> [!WARNING]
> Pushing a new Docker container affects the production dashboard deployed at NERSC through Spin.

> [!TIP]
> Run this workflow automatically with the Python script [publish_container.py](../publish_container.py):
> ```bash
> python publish_container.py --gui
> ```

> [!TIP]
> Prune old, unused images periodically in order to free up space on your machine:
> ```bash
> docker system prune -a
> ```

### Build the Docker image

1. Move to the root directory of the repository.

2. Build the Docker image:
```bash
docker build --platform linux/amd64 --output type=image,oci-mediatypes=true -t synapse-gui -f dashboard.Dockerfile .
```

### Push the Docker container

1. Move to the root directory of the repository.

2. Login to the [NERSC registry](https://registry.nersc.gov):
```bash
docker login registry.nersc.gov
# Username: your NERSC username
# Password: your NERSC password without 2FA
```

3. Tag the Docker image:
```bash
docker tag synapse-gui:latest registry.nersc.gov/m558/superfacility/synapse-gui:latest
docker tag synapse-gui:latest registry.nersc.gov/m558/superfacility/synapse-gui:$(date "+%y.%m")
```

4. Push the Docker container:
```bash
docker push -a registry.nersc.gov/m558/superfacility/synapse-gui
```

# References

* [Using NERSC's `registry.nersc.gov`](https://docs.nersc.gov/development/containers/registry/)
* [Superfacility API authentication](https://docs.nersc.gov/services/sfapi/authentication/#client)
The Synapse dashboard is documented in [docs/source/dashboard.md](../docs/source/dashboard.md).
20 changes: 20 additions & 0 deletions docs/Makefile
Original file line number Diff line number Diff line change
@@ -0,0 +1,20 @@
# Minimal makefile for Sphinx documentation
#

# You can set these variables from the command line, and also
# from the environment for the first two.
SPHINXOPTS ?=
SPHINXBUILD ?= sphinx-build
SOURCEDIR = source
BUILDDIR = build

# Put it first so that "make" without argument is like "make help".
help:
@$(SPHINXBUILD) -M help "$(SOURCEDIR)" "$(BUILDDIR)" $(SPHINXOPTS) $(O)

.PHONY: help Makefile

# Catch-all target: route all unknown targets to Sphinx using the new
# "make mode" option. $(O) is meant as a shortcut for $(SPHINXOPTS).
%: Makefile
@$(SPHINXBUILD) -M $@ "$(SOURCEDIR)" "$(BUILDDIR)" $(SPHINXOPTS) $(O)
12 changes: 12 additions & 0 deletions docs/docs.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,12 @@
name: synapse-docs

channels:
- conda-forge
- nodefaults

dependencies:
- myst-parser
- sphinx
- sphinx-autobuild
- sphinx-book-theme
- sphinx-copybutton
35 changes: 35 additions & 0 deletions docs/make.bat
Original file line number Diff line number Diff line change
@@ -0,0 +1,35 @@
@ECHO OFF

pushd %~dp0

REM Command file for Sphinx documentation

if "%SPHINXBUILD%" == "" (
set SPHINXBUILD=sphinx-build
)
set SOURCEDIR=source
set BUILDDIR=build

%SPHINXBUILD% >NUL 2>NUL
if errorlevel 9009 (
echo.
echo.The 'sphinx-build' command was not found. Make sure you have Sphinx
echo.installed, then set the SPHINXBUILD environment variable to point
echo.to the full path of the 'sphinx-build' executable. Alternatively you
echo.may add the Sphinx directory to PATH.
echo.
echo.If you don't have Sphinx installed, grab it from
echo.https://www.sphinx-doc.org/
exit /b 1
)

if "%1" == "" goto help

%SPHINXBUILD% -M %1 %SOURCEDIR% %BUILDDIR% %SPHINXOPTS% %O%
goto end

:help
%SPHINXBUILD% -M help %SOURCEDIR% %BUILDDIR% %SPHINXOPTS% %O%

:end
popd
30 changes: 30 additions & 0 deletions docs/source/conf.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,30 @@
# Configuration file for the Sphinx documentation builder.
#
# For the full list of built-in configuration values, see the documentation:
# https://www.sphinx-doc.org/en/master/usage/configuration.html

# -- Project information -----------------------------------------------------
# https://www.sphinx-doc.org/en/master/usage/configuration.html#project-information

project = "Synapse"
copyright = "BSD-3-Clause-LBNL"
author = "Arjun Dhamrait, Andrea Diaz, Marco Garten, Axel Huebl, Revathi Jambunathan, Remi Lehe, Ethan Rodriguez, Olga Shapoval, Jean-Luc Vay, Edoardo Zoni"

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

To discuss:

  • Names (so far, all contributors plus Jean-Luc Vay).
  • Order (so far, alphabetical with respect to the last name).


# -- General configuration ---------------------------------------------------
# https://www.sphinx-doc.org/en/master/usage/configuration.html#general-configuration

extensions = ["myst_parser", "sphinx_copybutton"]
myst_heading_anchors = 2

templates_path = ["_templates"]
exclude_patterns = []

# -- Options for HTML output -------------------------------------------------
# https://www.sphinx-doc.org/en/master/usage/configuration.html#options-for-html-output

html_theme = "sphinx_book_theme"
html_theme_options = {
"show_navbar_depth": 1,
"max_navbar_depth": 1,
}
html_static_path = []
Loading