Repository navigation
Understanding why MovingHorizonEstimator Hessians are always fully dense #368
Description
Activity
If this is due to the many zeros in your input data, SCT won't be able to do much (as you recall, we eventually reverted the changes inspired by #193). Have you tried storing
Aas aDiagonalfor instance, instead of an artificially dense matrix?Yes, the PR #216 introduced the changes to preserve
Diagonalmatrices for the weights. So the objective function is a sum of threedot(x, A, x), and theAargument is a sparseDiagonalmatrix for two of them.Now that I mention this, the first one,
dot(x̄, invP̄, x̄), is with a dense matrix, it may be the cause. I will investigate this avenue. Edit: btw,invP̄is really a dense (hermitian) matrix here, I cannot store it as aDiagonal.Edit 2: I'm referring to the three
dotcalls here:return dot(x̄, invP̄, x̄) + dot(Ŵ, invQ̂_Nk, Ŵ) + dot(V̂, invR̂_Nk, V̂) + Jε Oh sorry, the MRE example above is with a linear model, so it won't call this method of
obj_nonlinprog. My bad.Here's a better MRE with a
NonLinModel, to be coherent with my explanation above and maybe simplify the debugging (it leads to the same results as above):using ModelPredictiveControl, ControlSystemsBase Ts = 4.0 A = [ 0.800737 0.0 0.0 0.0 0.0 0.606531 0.0 0.0 0.0 0.0 0.8 0.0 0.0 0.0 0.0 0.6 ] Bu = [ 0.378599 0.378599 -0.291167 0.291167 0.0 0.0 0.0 0.0 ] Bd = [ 0; 0; 0.5; 0.5;; ] C = [ 1.0 0.0 0.684 0.0 0.0 1.0 0.0 -0.4736 ] Dd = [ 0.19; -0.148;; ] Du = zeros(2,2) f(x,u,d,p) = p.A*x + p.Bu*u + p.Bd*d h(x,d,p) = p.C*x + p.Dd*d model = NonLinModel(f, h, Ts, 2, 4, 2, 1, solver=nothing, p=(;A,Bu,Bd,C,Dd)) model = setop!(model, uop=[10,10], yop=[50,30], dop=[5]) using DifferentiationInterface, SparseConnectivityTracer, SparseMatrixColorings import ForwardDiff hessian = AutoSparse(AutoForwardDiff(), sparsity_detector=TracerSparsityDetector(), coloring_algorithm=GreedyColoringAlgorithm()) # hessian = AutoSparse(AutoForwardDiff(), sparsity_detector=TracerLocalSparsityDetector(), coloring_algorithm=GreedyColoringAlgorithm()) function gc!(LHS, X̂e, V̂e, Ŵe, Ue, Yem, De, P̄, x̄, p, ε) N = length(X̂e)÷6 LHS .= 0 for i in eachindex(LHS) if i ≤ N LHS[i] = X̂e[6*(i-1)+1] - 0 - ε end end return nothing end mhe = MovingHorizonEstimator(model; He=10, hessian, gc!, nc=11) using JuMP; unset_time_limit_sec(mhe.optim) res = sim!(mhe, 15, x̂_0=zeros(6), d_step=[-2.0]) info = getinfo(mhe) display(info[:∇²J])
Now that I mention this, the first one,
dot(x̄, invP̄, x̄), is with a dense matrix, it may be the cause. I will investigate this avenue. Edit: btw,invP̄is really a dense (hermitian) matrix here, I cannot store it as aDiagonal.So I just replaced:
return dot(x̄, invP̄, x̄) + dot(Ŵ, invQ̂_Nk, Ŵ) + dot(V̂, invR̂_Nk, V̂) + Jε
with:
return dot(Ŵ, invQ̂_Nk, Ŵ) + dot(V̂, invR̂_Nk, V̂)
and same issue. So this is not caused by the dense
invP̄matrix, nor theJεterm. And addingprintln(typeof(invR̂_Nk))andprintln(typeof(invQ̂_Nk))before thereturnreally shows that the two matrices are sparseLinearAlgebra.Hermitian{Float64, LinearAlgebra.Diagonal{Float64, Vector{Float64}}}. Note sure what is the cause...In which variables do the traced values live?
The decision vector
Z̃. It is defined as a concatenation of the slack, the arrival state estimate and the process noise over the time horizon:Z̃ = [ε; x̂0arr; Ŵ]. That is one of the goal of theupdate_prediction!call, to extractε,x̂0arrandŴfrom the decision vector, here:update_prediction!(x̂0arr, x̄, Ŵ, V̂, X̂0, Ŵe, V̂e, X̂e, û0, k, ŷ0, gc, g, estim, Z̃) Okay good news, it seems to be related to the computation of the estimated sensor noise
V̂fromx̂0arrandŴvectors, here:function predict_mhe!(
If I add aV̂.=0at the end of this function the pattern become sparse.The bad news is this function is really intricate and hard to debug.
Okay now I'm doubting that the zeros are in fact global zeros. I'm thinking that
SparseConnectivityTracermay be right and the pattern is indeed fully dense! I did not expect that, nor the Spanish Inquisition.For the linear case, we are able to analytically compute and inspect the Hessian. The Hessian is constant-in-time if the (inverted) arrival covariance
invP̄is constant (and if the data windows are filled). This can be achieved in this package by using aSteadyKalmanFilteras the arrival covariance estimator.The package computes the Hessian and its accessible in the field
H̃. Here's a generic linear example:using ControlSystemsBase, ModelPredictiveControl, LinearAlgebra, SparseArrays Ts = 4.0 A = [ 0.800737 1e-6 1e-6 1e-6 1e-6 0.606531 1e-6 1e-6 1e-6 1e-6 0.8 1e-6 1e-6 1e-6 1e-6 0.6 ] Bu = [ 0.378599 0.378599 -0.291167 0.291167 1e-6 1e-6 1e-6 1e-6 ] Bd = [ 0; 0; 0.5; 0.5;; ] C = [ 1.0 1e-6 0.684 1e-6 1e-6 1.0 1e-6 -0.4736 ] Dd = [ 0.19; -0.148;; ] Du = zeros(2,2) model = LinModel(ss(A,[Bu Bd],C,[Du Dd],Ts),Ts,i_d=[3]) model = setop!(model, uop=[10,10], yop=[50,30], dop=[5]) i_ym, nint_u, nint_ym, He = 1:2, [1, 1], 0, 5 Q̂ = diagm([1/2, 1, 1/2, 1, 1/2, 1].^2) R̂ = diagm([1, 1].^2) covestim = SteadyKalmanFilter(model, i_ym, nint_u, nint_ym, Q̂, R̂) P̂_0 = diagm([0.1, 0.1, 0.1, 0.1, 0.1, 0.1].^2) # user-specified fixed arrival covariance # set all off-diagonal coefficients of P̂_0 to 0.001: P̂_0 = P̂_0 .+ 0.001*ones(6, 6) - 0.001*I setstate!(covestim, zeros(6), P̂_0) mhe = MovingHorizonEstimator(model, He, i_ym, nint_u, nint_ym, P̂_0, Q̂, R̂; covestim) for i in 1:10 y = [2.0, 2.0] d = [1.0] x̂ = preparestate!(mhe, y, d) u = [0.0, 0.0] updatestate!(mhe, u, y, d) end display(sparse(mhe.H̃))
printing:
36×36 SparseMatrixCSC{Float64, Int64} with 1028 stored entries: ⎡⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⠀⣿⠀⎤ ⎢⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⠀⣿⠀⎥ ⎢⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⠀⣿⠀⎥ ⎢⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⠀⣿⠀⎥ ⎢⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⠀⣿⠀⎥ ⎢⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⠀⣿⠀⎥ ⎢⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⠀⣿⠀⎥ ⎢⠛⠛⠛⠛⠛⠛⠛⠛⠛⠛⠛⠛⠛⠛⠛⢄⠛⠀⎥ ⎣⠛⠛⠛⠛⠛⠛⠛⠛⠛⠛⠛⠛⠛⠛⠛⠀⠛⢄⎦The MHE is considered the analog of MPC but the more I study the MHE, the more I find differences compared to optimal control.
Sorry for the noise and thanks for your time once more. It seems that
SparseConnectivityTraceris smarter than me XD. Props to @adrhill.I will close the issue.
Reacted by Adrian Hill and Guillaume DalleHappy to hear SCT helped you gain some insights! :D
Reacted by Francis Gagnon and Guillaume DalleReacted by Francis Gagnon and Guillaume Dalle
Hello @gdalle !
I'm having the same problem as #193 but for the
MovingHorizonEstimator(MHE) instead of theNonLinMPC, that is, a fully-dense sparsity pattern for the Hessian of the objective function, when I'm 99.9% sure that many zeros are in fact global zeros. Note that I applied the equivalent improvements of #202 but for the MHE (on PR #216), so the problem seems to be elsewhere. I put too much time on trying to figure out why, so I'm asking for your help. What debugging method did you use to find out the two results mentioned here ? To be clearer, I want to handle structural sparsity better in the MHE, but I cannot find where the problem is inupdate_prediction!andobj_nonlinprogmethods.By using
ModelPredictiveControl.jlv2.4.1 (registered in a few minutes, else you candevit), here's a MRE:giving with
TracerSparsityDetector:and with
TracerLocalSparsityDetector:Many thanks for your time !