Per-row job metadata is useful precisely because _job_start_time / _job_duration / _job_version land on the row and outlive the job queue entry. But the documented way to read them back doesn't work — there's currently no public API that reaches them at all.
reference/specs/job-metadata.md (339-340) gives this as the access path:
# Access hidden attributes explicitly if needed:
SessionAnalysis().to_arrays('_job_start_time', '_job_duration', '_job_version')
That call raises.
Repro
import datajoint as dj
dj.config.jobs.add_job_metadata = True
dj.config.jobs.keep_completed = True
schema = dj.Schema('demo')
@schema
class Source(dj.Manual):
definition = """
source_id : int16
"""
@schema
class Result(dj.Computed):
definition = """
-> Source
---
value : float32
"""
def make(self, key):
self.insert1({**key, 'value': 1.0})
Source.insert1({'source_id': 1})
Result.populate()
The columns exist and are populated:
Result.connection.query(
f"SELECT _job_start_time, _job_duration, _job_version FROM {Result.full_table_name}"
).fetchall()
# ((datetime.datetime(2026, 9, 29, 21, 56, 26, 39000), 0.0117304, ''),)
The documented accessor doesn't reach them:
Result.to_arrays('_job_start_time')
# DataJointError: Attribute `_job_start_time` not found.
Cause
to_arrays projects before it fetches (src/datajoint/expression.py:1045):
projected = expr.proj(*fetch_attrs)
proj validates each requested name against heading.names (expression.py:572-575):
# check that all attributes exist in heading
try:
raise DataJointError("Attribute `%s` not found." % next(a for a in attributes if a not in self.heading.names))
except StopIteration:
pass # all ok
and names derives from the attributes property, which filters hidden attributes out (heading.py:263-272):
@property
def attributes(self):
...
# Excludes hidden attributes (names starting with ``_``).
return {k: v for k, v in self._attributes.items() if not v.is_hidden}
@property
def names(self) -> list[str]:
"""List of visible attribute names."""
return [k for k in self.attributes]
Nothing hidden survives proj, which closes off to_arrays, to_dicts, to_pandas and fetch together. Only Heading._attributes still sees them — that's how _has_job_metadata_attrs detects the columns (autopopulate.py:844) — and the writes go out as raw SQL (autopopulate.py:865).
The repo's own integration test works around it the same way, with a comment saying why (tests/integration/test_hidden_job_metadata.py:170-172):
# Fetch hidden attributes using raw SQL since fetch() filters them
result = conn.query(f"SELECT _job_start_time, _job_duration, _job_version FROM {table.full_table_name}").fetchall()
Conflicting guidance in the docs
reference/specs/table-declaration.md (223-238) describes the same filtering as intentional and directs readers to raw SQL:
# To inspect platform-managed hidden columns, query raw SQL.
# The public API (fetch / proj) intentionally rejects them.
The two pages give opposite answers, so one needs updating whichever way this is resolved. Flagging it here rather than separately, since the fix determines which one.
Why it matters
Writing metadata onto the row is what makes it outlive the queue entry, so reading it back is the point of the feature — how long a step took, which code version produced which rows, what was computed before a given commit. Raw SQL gets there, but it means interpolating full_table_name into a string and stepping outside the query layer: no restriction, no join, no fetch semantics, and backend-specific quoting by hand.
Not blocking — raw SQL is a workable stopgap. The ask is to not have to depend on it.
Environment
Confirmed on datajoint 2.3.3.dev16+g8e6ef3161 (current main), MySQL 8.0, Python 3.12. The block above is the run transcript, not a reconstruction. Line references are against main as of filing.
Per-row job metadata is useful precisely because
_job_start_time/_job_duration/_job_versionland on the row and outlive the job queue entry. But the documented way to read them back doesn't work — there's currently no public API that reaches them at all.reference/specs/job-metadata.md(339-340) gives this as the access path:That call raises.
Repro
The columns exist and are populated:
The documented accessor doesn't reach them:
Cause
to_arraysprojects before it fetches (src/datajoint/expression.py:1045):projvalidates each requested name againstheading.names(expression.py:572-575):and
namesderives from theattributesproperty, which filters hidden attributes out (heading.py:263-272):Nothing hidden survives
proj, which closes offto_arrays,to_dicts,to_pandasandfetchtogether. OnlyHeading._attributesstill sees them — that's how_has_job_metadata_attrsdetects the columns (autopopulate.py:844) — and the writes go out as raw SQL (autopopulate.py:865).The repo's own integration test works around it the same way, with a comment saying why (
tests/integration/test_hidden_job_metadata.py:170-172):Conflicting guidance in the docs
reference/specs/table-declaration.md(223-238) describes the same filtering as intentional and directs readers to raw SQL:The two pages give opposite answers, so one needs updating whichever way this is resolved. Flagging it here rather than separately, since the fix determines which one.
Why it matters
Writing metadata onto the row is what makes it outlive the queue entry, so reading it back is the point of the feature — how long a step took, which code version produced which rows, what was computed before a given commit. Raw SQL gets there, but it means interpolating
full_table_nameinto a string and stepping outside the query layer: no restriction, no join, no fetch semantics, and backend-specific quoting by hand.Not blocking — raw SQL is a workable stopgap. The ask is to not have to depend on it.
Environment
Confirmed on datajoint
2.3.3.dev16+g8e6ef3161(currentmain), MySQL 8.0, Python 3.12. The block above is the run transcript, not a reconstruction. Line references are againstmainas of filing.