Quick Start¶
This guide walks you through building and running your first BioImageFlow pipeline.
Files created by BioImageFlow¶
Before running a workflow, choose where BioImageFlow should write its state. There are three separate locations:
Workflow(storage_path=...)stores workflow outputs, cache metadata, and generated assets. If omitted, it defaults to./bif_data.configure_wetlands(wetlands_instance_path=...)stores Wetlands’ environment state: logs, debug port metadata, the bundled Pixi or Micromamba installation, and the isolated tool environments. Call it beforeWorkflow.compute(),Workflow.load(), orrequire_tool_packages().The tool store holds versioned tool packages installed by
require_tool_packages()orWorkflow.load(). It defaults to~/.bioimageflow/tool_packagesunlessBIOIMAGEFLOW_HOMEis set.
For a project-local run, start scripts like this:
from bioimageflow import Workflow, configure_wetlands
configure_wetlands(wetlands_instance_path="./wetlands")
with Workflow(storage_path="./bif_data") as wf:
...
If configure_wetlands() is not called, the Wetlands instance path
is resolved from BIOIMAGEFLOW_WETLANDS, then
BIOIMAGEFLOW_HOME / "wetlands", then ~/.bioimageflow/wetlands.
The tool store is resolved from BIOIMAGEFLOW_TOOL_STORE, then
BIOIMAGEFLOW_HOME / "tool_packages", then
~/.bioimageflow/tool_packages.
Pick a source¶
A workflow starts at a source node — a tool that produces the initial
DataFrame without consuming any upstream data. The companion
bioimageflow_common_tools package ships Files, which lists the files in
a directory:
from bioimageflow_common_tools import Files
files = Files()
Files is a DataFrameTool with
accepts_upstream = False. Its output is a path column containing
absolute paths. Source-tool patterns (and how to write your own) are covered
in Graph Construction.
Define a processing tool¶
A ProcessingTool runs on each row independently.
Declare its inputs, outputs, and the environment it needs:
from pathlib import Path
from typing import Annotated
from bioimageflow_core import (
ProcessingTool, GENERAL_ENV, ImageSpec, Arguments, Template,
)
class InvertImage(ProcessingTool):
display_name = "Invert"
environment = GENERAL_ENV
class Inputs:
image: Annotated[Path, ImageSpec()]
class Outputs:
inverted: Annotated[Path, ImageSpec()] = Template("{image.stem}_inv.tif")
def process_row(self, arguments: Arguments) -> "InvertImage.Outputs":
import imageio.v3 as iio
img = iio.imread(arguments.image)
out = 255 - img
iio.imwrite(arguments.inverted, out)
return self.Outputs(inverted=arguments.inverted)
Key points:
Inputsfields that receive image data useAnnotated[Path, ImageSpec(...)].Outputspath fields useTemplate(...)defaults for output path templates (see Output Templating).process_rowreceives anArgumentsobject with resolved values for every input and output field.
Build the workflow¶
Wire tools together inside a Workflow context manager:
from bioimageflow import Workflow, configure_wetlands
invert = InvertImage()
configure_wetlands(wetlands_instance_path="./wetlands")
with Workflow(storage_path="./bif_data") as wf:
images = files(path="/data/raw", pattern="*.tif")
inverted = invert(image=images["path"], name="invert")
result = wf.compute(inverted)
print(result)
# inverted
# 0 /.../bif_data/cache/v1/results/.../records/rec_.../assets/image1_inv.tif
# 1 /.../bif_data/cache/v1/results/.../records/rec_.../assets/image2_inv.tif
What happens:
files(path=...)creates a graph node — no computation yet.invert(image=images["path"])binds theimageinput to thepathcolumn of the source node.wf.compute(inverted)executes the DAG in topological order.
The result is a pandas DataFrame with one column per output field.
Owned output assets are stored in immutable records under
./bif_data/cache/v1/results/.../<result-key>/records/<record-id>/ and the
JSON run view under ./bif_data/views/runs/ points back to the selected record.
Pass output_view="symlink" or output_view="copy" to Workflow when
you also want browseable files under ./bif_data/outputs/latest/.
Caching and Provenance covers the cache lifecycle, plan() and
invalidate(). Wetlands environments for this run are kept under
./wetlands because the script called configure_wetlands().
Re-running is free¶
Run the same workflow again and it completes instantly — the cache recognises that input references, parameters, and tool versions haven’t changed:
with Workflow(storage_path="./bif_data") as wf:
images = files(path="/data/raw", pattern="*.tif")
inverted = invert(image=images["path"], name="invert")
result = wf.compute(inverted) # cache hit, no recomputation
Next steps¶
Graph Construction — nodes, column references, branching, and source tools
Writing Custom Tools — writing your own ProcessingTool and DataFrameTool
Environments — Wetlands paths, environment configuration, and per-environment runtime settings
Architecture — understand the two-package design