Computational Models
A model with a spatial component predicts the response to stimulation by a
particular device, so it is bound to an implant and handed the stimulus.
A temporal-only model describes one location’s response over time and needs
no implant. Most users work with
Model, which can contain a spatial component,
a temporal component, or both:
SpatialModeldetermines where stimulation appears in the visual field.TemporalModeldetermines how the response evolves over time.
Available models
Reference |
Model |
Type |
|---|---|---|
generic |
temporal |
|
generic |
temporal |
|
spatial |
||
temporal |
||
spatial + temporal |
||
spatial |
||
spatial |
||
spatiotemporal |
||
spatiotemporal |
Cortical stimulation also has
ScoreboardModel, a spatial baseline
that maps cortical electrode locations through cortical retinotopy.
Which model to use depends on the scientific question. The scoreboard model is a simple local baseline. The axon-map model adds retinal nerve-fiber effects for epiretinal stimulation. Published temporal and spatiotemporal models add assumptions specific to their experiments and should be chosen when those assumptions are relevant.
Basic usage
Models follow the same workflow: choose an implant, bind a model to it, then predict a percept from a stimulus.
import pulse2percept as p2p
implant = p2p.implants.ArgusII()
model = p2p.models.ScoreboardModel(implant=implant, rho=200)
percept = model.predict_percept({'A8': 30})
The result of predict_percept is a
Percept.
Source, delivered stimulation, percept
The prediction pipeline distinguishes the source, delivered stimulation, and percept:
source → implant → delivered stimulation → model → percept
- Source
Input presented to the device: a
Stimulus(or compatible scalar, array, or dict),ImageStimulus,VideoStimulus, orScene.- Delivered stimulation
Electrical stimulation after implant preprocessing, encoding, raster scheduling, threshold calibration, and safety checks. Models call
implant.prepare_stim(source)internally; call it directly to inspect the delivered stimulus.- Percept
Model output from
model.predict_percept(source).
Building
Models build automatically on first prediction. Changing a model parameter invalidates the affected component, which is rebuilt when needed:
model = p2p.models.AxonMapModel(implant=implant)
# Builds automatically:
percept = model.predict_percept(stim)
# Rebuilds the spatial component:
model.spatial.rho = 250
percept = model.predict_percept(stim)
Rebinding the implant also invalidates the spatial build because it depends on
device geometry. model.build() forces a full rebuild; to set parameters as
you build a component, use model.spatial.build(rho=250).
Electrode-retina distance
ScoreboardModel,
AxonMapModel and
Thompson2003Model use electrode x and
y coordinates only. Nonzero z values therefore do not affect their
output and produce a warning.
This is a model limitation. Electrode-target distance is expected to affect
stimulation threshold and spatial recruitment, but pulse2percept does not
currently parameterize that relationship because the required psychophysical
evidence is insufficient. In the Scoreboard and AxonMap models, rho remains
an effective perceptual spread parameter fitted to subject reports rather than
inferred from electrode-retina distance.
Simulating a visual scene
Added in version 0.11.0.
The workflow above starts from a stimulus you built yourself. To start from
what someone is looking at instead, give the model a
Scene:
from pulse2percept.units import dva
scene = p2p.vision.Scene(p2p.stimuli.LogoBVL(), fov=40 * dva)
implant = p2p.implants.ArgusII()
implant.encoder = p2p.stimuli.AmplitudeEncoder(amp_range=(0, 50))
model = p2p.models.ScoreboardModel(implant=implant, rho=200)
percept = model.predict_percept(scene, gaze=(0, 0) * dva)
Scene prediction separates four responsibilities:
Scene |
What is visually present, and where native vision is lost. |
Implant |
Device geometry and encoding constraints. |
Model |
Knows the retinotopy, and so connects Scene to Implant. |
Percept |
What the simulated observer sees. |
The model maps implant coordinates into the visual field through its retinotopic map. Each electrode follows this chain:
retinal coordinate (um)
-> visual_field_map.ret_to_dva -> eye-centered visual field (dva)
-> + gaze, for eye-coupled input only -> scene coordinate (dva)
-> sample the scene
gaze is the scene location that currently falls on the fovea, so
scene = eye-centered visual field + gaze. Gaze always decides where the
percept lands in scene coordinates. Whether it also decides what the
electrodes are given depends on the implant’s
scene_input_frame:
|
Input passes through the eye’s optics (Alpha, PRIMA), so gaze moves the scene across the implant as well as moving the percept across the scene. This is the default. |
|
Input comes from a head-fixed camera (Argus, BVT, IMIE) that the eye cannot move: the electrodes are handed the same scene whatever the gaze, and only the percept moves. |
Neither the implant nor an eye-centered
Scotoma moves when gaze does. Pass one
(x, y) to fixate, or one per video frame to move the eye between frames.
For a 'head' system, the sampling locations above are still the electrodes’
own visual-field positions, which assumes the device’s camera-to-electrode
registration is aligned with them. Real systems configure that mapping
separately and it is not modeled here.
scene_input_frame follows the device class but is a property of the
system, so one implant can be run the other way – an Argus II with eye
tracking, which shifts the camera ROI with gaze:
implant = p2p.implants.ArgusII()
implant.scene_input_frame = 'eye'
The sampled values are passed to implant.encoder, which maps gray levels
to current and applies device timing constraints. A scene is per-prediction
input and is not stored on the model or implant.
An implant’s preprocess – an edge filter, an inversion, a contrast
stretch – is applied to the prosthetic input branch only, before the
scene is sampled at the electrode locations, because an image operation needs
an image and by sampling time there is one number per electrode. Native and
residual vision always use the original scene: what the device does to its
own input is not something the eye goes through. Spatial preprocessing
operates at the scene source’s pixel resolution.
implant.preprocess = lambda stim: stim.filter('sobel')
For scene input, preprocess must return an
ImageStimulus or
VideoStimulus; conversion to electrical
stimulation belongs to the encoder. Pixel values and channels may change, but
spatial shape and frame timing must remain unchanged because fov and the
frame clock refer to the original scene.
Scene registration currently requires a retinal visual_field_map and an
implant encoder. A cortical visual_field_map or a missing encoder
raises ValueError.
Residual vision
If the scene also carries a Scotoma, the
result is what the person actually sees – intact native vision outside the
lost region, and the prosthetic percept inside it – as a single RGB
Percept on the scene’s own pixel grid:
scene = p2p.vision.Scene(p2p.stimuli.LogoBVL(), fov=40 * dva,
scotoma=p2p.vision.Scotoma.circle(8 * dva))
model = p2p.models.ScoreboardModel(implant=implant, rho=200)
percept = model.predict_percept(scene, gaze=(0, 0) * dva, vmax=50)
vmax is required here and is not inferred: model brightness is in
arbitrary units, so which brightness counts as white is a claim about the
display, not about the model. Holding it fixed across calls is what keeps two
gazes comparable.
The scotoma affects native vision only. Prosthetic encoding samples the unmasked scene, including locations inside the scotoma.
Percept data layouts
A Percept holds one of two layouts, with
time as the last axis in both:
(Y, X, T) perceived brightness in arbitrary units
(Y, X, 3, T) RGB intensities in [0, 1]
Prosthesis models produce brightness percepts. When a
Scene has a scotoma, scene-driven prediction
composes that model output with residual vision and returns an RGB percept:
import numpy as np
from pulse2percept.percepts import Percept
rgb = Percept(np.zeros((60, 80, 3, 1)))
rgb.is_rgb # True
rgb[..., 0].shape # (60, 80, 3): one frame, still in color
rgb.plot() # drawn as RGB, without a colormap
RGB values are display intensities and must be finite and lie in [0, 1];
anything else raises at construction rather than saturating quietly later. The
RGB axis is not a spatial dimension: space still describes (Y, X).
Operations defined on perceived brightness – n_gray, noise,
argmax, max, and the vmin/vmax display range – raise a
ValueError for an RGB percept rather than inventing a conversion from
color to brightness. Ranking three channels by one number would have to pick a
color metric, which is also why a multi-frame RGB percept has no brightest
frame to plot(); animate it with play() instead. percept.data is
always available for the plain numerical answer.
Spatial and temporal components
Classes ending in Model are complete models with explicit constructor
parameters:
model = p2p.models.AxonMapModel(
implant,
rho=300,
lam=500,
)
percept = model.predict_percept(stim)
Classes ending in Spatial or Temporal are components, and
Model combines two of them:
spatial = p2p.models.AxonMapSpatial(
implant,
rho=300,
lam=500,
)
temporal = p2p.models.FadingTemporal(tau=100)
model = p2p.models.Model(spatial, temporal)
Use Model to combine spatial and temporal components from different
models. At least one component is required, and each must already be
constructed. The implant belongs to the spatial component.
Parameters
Component parameters are accessed directly:
model.spatial.rho = 250
model.temporal.tau = 50
Named-model constructors expose the same parameters directly. After
construction, access them through the component. Parameters declared by both
components, such as thresh_percept, remain independent.
The API reference for each model documents its assumptions, parameters, input requirements, and numerical units.