Utilities for loading and dumping database data as JSON.
These utilities (partially) replace Django's built-in dumpdata and
loaddata management commands.
Suppose you want to move data between systems incrementally. In this case it
isn't sufficient to only know the data which has been created or updated; you
also want to know which data has been deleted in the meantime. Django's
dumpdata and loaddata management commands only support the former case,
not the latter. They also do not including dependent objects in the dump.
This package offers utilities and management commands to address these shortcomings.
pip install feincms3-data.
Add feincms3_data to INSTALLED_APPS so that the included management
commands are discovered.
Add datasets somewhere describing the models and relationships you want to
dump, e.g. in a module named app.f3datasets:
from feincms3_data.data import (
specs_for_app_models,
specs_for_derived_models,
specs_for_models,
)
from app.dashboard import models as dashboard_models
from app.world import models as world_models
def districts(args):
pks = [int(arg) for arg in args.split(",") if arg]
return [
*specs_for_models(
[world_models.District],
{
"filter": {"pk__in": pks},
"delete_missing": True,
},
),
*specs_for_models(
[world_models.Exercise],
{
"filter": {"district__in": pks},
"delete_missing": True,
},
),
# All derived non-abstract models which aren't proxies:
*specs_for_derived_models(
world_models.ExercisePlugin,
{
"filter": {"parent__district__in": pks},
"delete_missing": True,
},
),
]
def datasets():
return {
"articles": {
"specs": lambda args: specs_for_app_models(
"articles",
{"delete_missing": True},
),
},
"pages": {
"specs": lambda args: specs_for_app_models(
"pages",
{"delete_missing": True},
),
},
"teachingmaterials": {
"specs": lambda args: specs_for_models(
[
dashboard_models.TeachingMaterialGroup,
dashboard_models.TeachingMaterial,
],
{"delete_missing": True},
),
},
"districts": {
"specs": districts,
},
}Add a setting with the Python module path to the specs function:
FEINCMS3_DATA_DATASETS = "app.f3datasets.datasets"Now, to dump e.g. pages you would run:
./manage.py f3dumpdata pages > tmp/pages.json
To dump the districts with the primary key of 42 and 43 you would run:
./manage.py f3dumpdata districts:42,43 > tmp/districts.json
The resulting JSON file has three top-level keys:
"version": 1: The version of the dump, because not versioning dumps is a recipe for pain down the road."specs": [...]: A list of model specs."objects": [...]: A list of model instances; uses the same serializer as Django'sdumpdata, everything looks the same.
Model specs consist of the following fields:
"model": The lowercased label (app_label.model_name) of a model."filter": A dictionary which can be passed to the.filter()queryset method as keyword arguments; used for determining the objects to dump and the objects to remove after loading."delete_missing": This flag makes the loader delete all objects matching"filter"which do not exist in the dump. Those objects whose deletion is a precondition for loading the dump at all -- because they hold unique values which an object from the dump is claiming -- are deleted before loading instead of at the end. Nothing else changes: the very same objects are deleted, only earlier."ignore_missing_m2m": A list of field names where deletions of related models should be ignored when restoring. This may be especially useful when only transferring content partially between databases."save_as_new": If present and truish, objects are inserted using new primary keys into the database instead of (potentially) overwriting pre-existing objects."defer_values": A list of fields which should receive random garbage when loading initially and only receive their real value later. This is especially useful to avoid unique constraint errors when loading partial graphs.
Note
Multi table inheritance children share the primary key of their parent. If
the database says an object is a different concrete model than the dump does
-- which happens once the databases drift apart, e.g. because objects are
created on the target as well -- loading is refused with an
InconsistentModelError.
Loading anyway would leave the stale row of the other type behind: the parent row is shared, so nothing ever removes it, and the result would be an object which is two things at once. Removing it automatically isn't an option either -- deleting the stale child takes the shared parent row with it, and since Django is perfectly happy with a parent having several children the row may not even be stale. Delete the offending objects yourself and load again.
Note
Objects which have been deleted and recreated on the source database arrive with a new primary key, while the target database still holds the row with the same unique values. Databases don't allow both rows to exist at the same time, so the old row has to go before the dump can be loaded.
"delete_missing" handles this by itself, as long as the old row matches
the spec's "filter". If you cannot use "delete_missing" for a model
-- typically because deletions shouldn't be propagated to the target as soon
as anything else is transferred -- restrict the filter to the unique values
contained in the dump instead:
specs_for_models(
[Identifier],
{
"filter": {"identifier__in": identifiers},
"delete_missing": True,
},
)This only ever deletes rows claiming one of the dumped identifiers (and everything hanging off them) and leaves all other identifiers alone. Keep the filter in sync with the objects you're actually dumping.
Note
When using save_as_new and delete_missing together, you may need to
specify how primary keys should be mapped to avoid inadvertent deletion of
objects. If you have a parent model with save_as_new and child models
with both save_as_new and delete_missing, you should use the
dictionary form of delete_missing with a map parameter to map old
primary keys to new ones. For example:
{
"filter": {"parent__in": [42]},
"save_as_new": True,
"delete_missing": {
"map": [
("parent__in", "app.Parent"),
],
},
}This ensures that child objects are matched against the new parent primary keys rather than the old ones, preventing old data from being kept and new data from being inadvertently deleted.
Note that the mapping fails loudly on purpose if the primary key mapping
isn't available. The reason could be that the parent model doesn't actually
use save_as_new. Since this is a potentially destructive operation it's
better to fail loudly than to silently eat data.
The dumps can be loaded back into the database by running:
./manage.py f3loaddata -v2 tmp/pages.json tmp/districts.json
Each dump is processed in an individual transaction. The data is first loaded
into the database; at the end, data matching the filters but whose primary
key wasn't contained in the dump is deleted from the database (if
"delete_missing": True). The only exception are objects holding unique
values which the dump's data claims -- those are removed upfront, since
databases do not allow the old and the new row to exist at the same time.
Both deletions are restricted to the spec's "filter". An object outside of
the filter which holds a unique value claimed by the dump therefore still makes
the load fail; widen the filter (or dump fewer objects) in that case.