Skip to content

Working with expensive notebooks

marimo provides tools to control when cells run. Use these tools to prevent expensive cells, which may call APIs or take a long time to run, from accidentally running.

Stop execution with mo.stop

Use mo.stop to stop a cell from executing if a condition is met:

# if condition is True, the cell will stop executing after mo.stop() returns
mo.stop(condition)
# this won't be called if condition is True
expensive_function_call()

Use mo.stop with mo.ui.run_button() to require a button press for expensive cells:

Configure how marimo runs cells

Disable cell autorun

If you habitually work with very expensive notebooks, you can disable automatic execution. When automatic execution is disabled, when you run a cell, marimo marks dependent cells as stale instead of running them automatically.

Disable autorun on startup

marimo autoruns notebooks on startup, with marimo edit notebook.py behaving analogously to python notebook.py. This can also be disabled through the notebook settings.

Disable individual cells

marimo lets you temporarily disable cells from automatically running. This is helpful when you want to edit one part of a notebook without triggering execution of other parts. See the reactivity guide for more info.

Manage memory

Here are a few tips for managing the memory consumption of your notebooks, on host or GPU.

Wrap intermediate computations in functions

By default, global variables live in the kernel memory. Intermediate variables that are defined in functions are cleaned up automatically.

For example, if X is a temporary:

Do this:

def _():
    X = torch.randn(1e4, 1e4, device='cuda')
    Y = f(X)
    return Y

Don't do this:

X = torch.randn(1e4, 1e4, device='cuda')
Y = f(X)
# X still lives in program memory!

Use del to remove variables from kernel memory

Use the del operator to remove variables from kernel memory.

In a single cell. Prefer deleting variables in the cell they were defined in. For example, if X is a temporary that you don't need after computing Y:

X = torch.randn(1e4, 1e4, device='cuda')
Y = f(X)
del X

In another cell. Sometimes, computations are spread across multiple cells, and you only realize later on that you need to free memory that you've already allocated. In such cases you can still use the del keyword. For example:

data = load_large_dataset()
derived_data = f(data)
del data

marimo inserts control dependences to make sure that variables are not deleted before they are used. When del is used to delete a variable that was defined in a another cell, the cell where del was used becomes a child of all other cells that reference that variable. In this case, that means marimo knows to run the third cell after the second cell, since the second cell references data and the third cell deletes it. However, once data is deleted, attempting to manually run the second cell will raise a NameError, and you'll need to re-run the defining cell in order to get your notebook back to a consistent state.

Local variables

Local or temporary variables (i.e., variables prefixed with an underscore) are automatically removed from the kernel globals after cell-run. If another Python object retains a reference to the variable, it will remain in memory; otherwise, Python's garbage collector will automatically reclaim its allocated memory.

Automatically snapshot outputs as HTML or IPYNB

To keep a record of your cell outputs while working on your notebook, you can configure notebooks to automatically save as HTML or ipynb through the notebook menu (these files are saved in addition to the notebook's .py file). Snapshots are saved to a folder called __marimo__ in the notebook directory.

Learn more about exporting notebooks in our exporting guide.

Cache expensive computations

marimo provides two decorators to cache the return values of expensive functions:

  1. In-memory caching with mo.cache
  2. Disk caching with mo.persistent_cache

Both utilities can be used as decorators or context managers.

import marimo as mo

@mo.cache
def compute_embedding(data: str, embedding_dimension: int, model: str) -> np.ndarray:
    ...
import marimo as mo

@mo.persistent_cache
def compute_embedding(data: str, embedding_dimension: int, model: str) -> np.ndarray
    ...

See our guide on caching for details, including how the cache key is constructed, and limitations.

Tip

Use the marimo cache command group to see and reclaim the disk space used by mo.persistent_cache.

marimo cache dir notebook.py

Run marimo cache dir ./dir to print cache directories resolved from the given folder.

Check disk usage

marimo cache size -r notebooks/

Run marimo cache size to print the disk usage and entry count for each cache directory, plus a total when there is more than one.

Delete entries outright

Each mo.persistent_cache name gets its own block, the subdirectory that holds its entries, for example train in __marimo__/cache/train/. Anything in the name other than letters, digits, spaces, _, and - is replaced with _ in the block's name, so two names that differ only there share a block. A cache directory can also hold a manifest, a file that records the cache keys a notebook produced.

marimo cache clean my_notebook.py train

Run marimo cache clean to delete cache entries. With a notebook PATH, it deletes exactly the entries listed in that notebook's manifest, then empties those records from the manifest. Entries the manifest does not list are left in place. With a directory PATH, it deletes whole blocks, including the blob directories that hold large entry values. If you name no blocks, it deletes every block. Pass one or more NAME arguments to limit either mode to those blocks. Each NAME must match a name you gave to mo.persistent_cache.

Delete entries that current code can no longer produce

marimo cache prune my_notebook.py --dry-run

Run marimo cache prune to delete cache entries that current code can no longer produce, based on the manifest. For example, if you edit a cell that a cached block depends on, prune treats that block's older entries as dead and deletes them. The next run of the notebook records fresh entries under the new code.

Prune never deletes an entry recorded under code that still exists. Deleting too much costs a recomputation, never a wrong result.

--dry-run reports the planned deletions and deletes nothing. Remove the flag to delete the entries. Add --force (or -y) to skip the confirmation prompt, for example in an automated script.

marimo cache prune my_notebook.py --force

See Managing the cache directory for the manifest, the path hash, and the guards and caveats behind prune.

Lazy-load expensive UIs

Lazily render UI elements that are expensive to compute using marimo.lazy.

For example,

import marimo as mo

data = db.query("SELECT * FROM data")
mo.lazy(mo.ui.table(data))

In this example, mo.ui.table(data) will not be rendered on the frontend until is it in the viewport. For example, an element can be out of the viewport due to scroll, inside a tab that is not selected, or inside an accordion that is not open.

However, in this example, data is eagerly computed, while only the rendering of the table is lazy. It is possible to lazily compute the data as well: see the next example.

import marimo as mo

def expensive_component():
    import time
    time.sleep(1)
    data = db.query("SELECT * FROM data")
    return mo.ui.table(data)

accordion = mo.accordion({
    "Charts": mo.lazy(expensive_component)
})

In this example, we pass a function to mo.lazy instead of a component. This function will only be called when the user opens the accordion. In this way, expensive_component lazily computed and we only query the database when the user needs to see the data. This can be useful when the data is expensive to compute and the user may not need to see it immediately.