Python-Related Codes and Scripts
Python virtual environment.
This is to run on Genoa CPU nodes of Owl.
# To make the VE on Owl
# ---------------------
# You will need to use your own account.
interact --account=arcadm --time=2:00:00 --nodes=1 --tasks-per-node=1 --cpus-per-task=1 --partition=normal_q --constraint=avx512
module reset
module load Miniforge3/25.11.0-1
## You will want to choose your own path: this is a path on my machine.
## I use these paths to structure where I put VEs.
conda create -p /home/ckuhlman/env-python/owl/normal_q/genoa/py314_mf_statsmodels
## Alter the path here similarly.
source activate /home/ckuhlman/env-python/owl/normal_q/genoa/py314_mf_statsmodels
conda install python=3.14
conda install pandas
conda install matplotlib
conda install -c conda-forge statsmodels
# If ipykernel does not show up in "conda list", then:
pip install ipykernel
python -m ipykernel install --user --name py314_mf_statsmodels --display-name "Linear StatsModels py3.14 owl genoa"
# All packages installed.
# Deactivate the VE.
conda deactivate
# Now done with making the virtual environment (VE).
# We are still on a compute node.
# =================================================
# The next time you want to use the VE, just use:
module load Miniforge3/25.11.0-1
# This path must be the same as those above.
source activate /home/ckuhlman/env-python/owl/normal_q/genoa/py314_mf_statsmodels
Python code.
File linear-fit.py:
# Purpose: demonstrate simple plotting of x-y data and linear fit.
# Modules:
# module load
import time
import sys
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
import statsmodels.formula.api as smf
import statsmodels.api as sma
begin_time=time.time()
# CLAs (command line arguments).
input_filename=sys.argv[1]
output_filename=sys.argv[2]
x_lower_fit=(int)(sys.argv[3])
x_upper_fit=(int)(sys.argv[4])
# Load the data into a df.
df = pd.read_csv(input_filename, sep=' ')
# Print the data to confirm.
print(" Contents of dataframe (df):")
print(df)
# Fit the ordinary least squares (OLS) model
# The formula 'y_values ~ x_values' automatically includes an intercept (b)
model = smf.ols(formula='response ~ predictor', data=df)
results = model.fit()
# Print out only the calculated coefficients (Intercept and Slope)
print("Intercept (b):", results.params['Intercept'])
print("Slope (m):", results.params['predictor'])
print(" model summary: ")
print(results.summary())
# Generate the line from the least squares fit.
x_fit=list()
y_fit=list()
x_fit.append(x_lower_fit)
x_fit.append(x_upper_fit)
y_val = results.params['predictor'] * x_lower_fit + results.params['Intercept']
y_fit.append(y_val)
y_val = results.params['predictor'] * x_upper_fit + results.params['Intercept']
y_fit.append(y_val)
# Plot the results.
# Plot the original data points and the fit line
plt.scatter(df.predictor, df.response, color='blue', alpha=0.6, label='Data Points')
plt.plot(x_fit, y_fit, color='red', linewidth=2, label='Statsmodels Fit')
plt.xlabel('predictor')
plt.ylabel('response')
plt.legend()
plt.savefig(output_filename)
end_time=time.time()
duration = end_time - begin_time
print(" execution duration (s) : ",duration)
Sbatch script for batch job on Owl.
File sbatch.python.owl.genoa.slurm:
#!/bin/bash
## Python on Owl.
## -----------------------
## ACCOUNT.
## The account to charge to.
## You will have your own accounts.
#SBATCH --account arcadm
## -----------------------
# SLURM JOB SCRIPT OPTIONS:
#SBATCH --job py-lin-fit
## -----------------------
## EXECUTION DURATION.
# Set the time, which is the maximum time your job can run in HH:MM:SS.
#SBATCH --time=0:10:00
## -----------------------
## NUM NODES AND CORES.
# A serial code needs 1 node, 1 task, and 1 cpu.
## Number of tasks.
#SBATCH --ntasks=1
## Number of tasks per node (can compute number of nodes).
#SBATCH --ntasks-per-node=1
## Number of cores (total) per task.
#SBATCH --cpus-per-task=1
## -----------------------
## JOB QUEUE/PARTITION AND CONSTRAINTS.
## Set the partition to submit to (a partition is equivalent to a queue)
#SBATCH --partition=normal_q
#SBATCH --constraint=avx512
## -----------------------
## MEMORY.
## If needed, request whole nodes by uncommenting the "#SBATCH --exclusive" line
## below.
## #SBATCH --exclusive
##SBATCH --mem=122G
## -----------------------
## SLURM OUTPUT AND ERROR FILES.
#SBATCH --output slurm.owl.genoa.python.linear.fit.%j.out
#SBATCH --error slurm.owl.genoa.python.linear.fit.%j.err
## -----------------------
## RESERVATION.
## #SBATCH --reservation=HPCMaintenance
## -----------------------
## MODULES AND VES.
module reset
module load Miniforge3/25.11.0-1
## You will have to change the path and name to the VE you created.
source activate /home/ckuhlman/env-python/owl/normal_q/genoa/py314_mf_statsmodels
## -----------------------
## WORKING DIRECTORY.
cd $SLURM_SUBMIT_DIR
## -----------------------
## Record slurm conditions.
echo " "
echo "slurm scontrol:"
echo " "
scontrol show job --details $SLURM_JOB_ID
echo " "
echo " << end slurm scontrol >>"
echo " "
## -----------------------
## EXPORTS
## Exports and variable assignments.
export OMP_NUM_THREADS=$SLURM_CPUS_PER_TASK
export MV2_ENABLE_AFFINITY=0
echo " SLURM_CPUS_PER_TASK: " $SLURM_CPUS_PER_TASK
echo " OMP_NUM_THREADS: " $OMP_NUM_THREADS
echo " MV2_ENABLE_AFFINITY: " $MV2_ENABLE_AFFINITY
echo " SLURM_NTASKS: " $SLURM_NTASKS
echo " SLURM_JOB_NUM_NODES: " $SLURM_JOB_NUM_NODES
## -----------------------
## JOB.
sh run.me
Bash script to launch Python code from slurm script.
File run.me:
x_lower_fit=40
x_upper_fit=100
code=linear-fit.py
input_filename=in_data.inp
# output_filename=python.plot.out.pdf
output_filename=python.plot.out.png
python ${code} ${input_filename} ${output_filename} ${x_lower_fit} ${x_upper_fit}