[indexer-alt] Add pruner pipeline for obj_info #20539

lxfind · 2024-12-06T22:23:18Z

Description

This PR implements the obj_info_pruner pipeline.
I refactored the obj_info pipeline so that these two pipelines could share the same process function logic.
For obj_info_pruner, it look at every (obj_id, checkpoint) pair produced from the obj_info processing, and prune accordingly. Within each commit, it does sequential pruning which is less ideal, but hopefully if we have enough concurrent at the outer layer, this is not a huge problem.

Test plan

How did you test the new or updated feature?

Release notes

Check each box that your changes affect. If none of the boxes relate to your changes, release notes aren't required.

For each box you select, include information after the relevant heading that describes the impact of your changes that a user might notice and any actions they must take to implement updates.

vercel · 2024-12-06T22:23:22Z

The latest updates on your projects. Learn more about Vercel for Git ↗︎

Name	Status	Preview	Comments	Updated (UTC)
sui-docs	✅ Ready (Inspect)	Visit Preview	💬 Add feedback	Dec 10, 2024 1:50am

3 Skipped Deployments

Name	Status	Preview	Updated (UTC)
multisig-toolkit	⬜️ Ignored (Inspect)	Visit Preview	Dec 10, 2024 1:50am
sui-kiosk	⬜️ Ignored (Inspect)	Visit Preview	Dec 10, 2024 1:50am
sui-typescript-docs	⬜️ Ignored (Inspect)	Visit Preview	Dec 10, 2024 1:50am

amnn

Some thoughts/questions on consistency between impls but mostly I'm curious/scared about issuing each delete as its own request to the DB 😬

amnn · 2024-12-09T15:45:17Z

crates/sui-indexer-alt/src/config.rs

@@ -246,7 +248,7 @@ impl ConcurrentLayer {
                (None, _) | (_, None) => None,
                (Some(pruner), Some(base)) => Some(pruner.finish(base)),
            },
-            checkpoint_lag: self.checkpoint_lag.or(base.checkpoint_lag),
+            checkpoint_lag: base.checkpoint_lag,


It would be nice to keep the logic here consistent with the sequential layer (although I don't particularly mind whether we permit overriding per pipeline or not) -- is there a reason we can't do that?

Because ConcurrentLayer does not have the checkpoint_lag field.

But you added checkpoint_lag to ConcurrentConfig right? Why not add it to ConcurrentLayer? Even if it's unlikely that someone would want to set a checkpoint lag for all concurrent layers, it seems confusing to selectively offer the override logic -- it breaks the mental model for people who are interacting with this mainly be reading the configs.

An alternative is to remove this field entirely even from ConcurrentConfig.
The primary intention is indeed to make sure that users cannot specify a global lag for all concurrent pipelines, since it does not make sense. This field comes from the consistency layer, instead of concurrent layer.

Ended up adding it to ConcurrentLayer. I realize this is not part of the global indexer config, so it's probably fine.

amnn · 2024-12-09T15:58:22Z

crates/sui-indexer-alt/src/handlers/obj_info_pruner.rs

+        // TODO: We could consider make this more efficient by doing some grouping in the collector
+        // so that we could merge as many objects as possible across checkpoints.


We can maybe make this slightly better by setting a long commit interval? That way each individual delete might clean up multiple old versions of a given object?

crates/sui-indexer-alt/src/handlers/obj_info_pruner.rs

amnn · 2024-12-09T16:02:46Z

crates/sui-indexer-alt/src/handlers/obj_info_pruner.rs

+        });
+        let mut committed_rows = 0;
+        for (object_id, cp_sequence_number_exclusive) in to_prune {
+            committed_rows += diesel::delete(obj_info::table)


I'm ...pretty nervous about this.

amnn · 2024-12-09T16:24:34Z

crates/sui-indexer-alt/src/handlers/obj_info.rs

+pub(crate) enum ProcessedObjInfoUpdate {
+    Insert(Object),
+    Delete(ObjectID),
+}
+
+pub(crate) struct ProcessedObjInfo {
+    pub cp_sequence_number: u64,
+    pub update: ProcessedObjInfoUpdate,
+}


nit: Could you move the top-level elements around for consistent file order? types at the top, then impl blocks, then trait impls, then free functions.

nit (optional): I think I started with something quite similar to this pattern for StoredObjectUpdate, but ended up going for something more like this (not using a special enum type and pulling the object_id into the outer struct) as it made some things neater:

pub(crate) struct ProcessedObjInfo { pub object_id: ObjectID, pub cp_sequence_number: u64, pub update: Option<Object> }

nit: Could you move the top-level elements around for consistent file order? types at the top, then impl blocks, then trait impls, then free functions.

hmm isn't it always better for impls to be right next to its struct definition? Otherwise it makes it diffult to locate them.

This ordering of elements in the file helps in two ways:

It's consistent across the codebase (in this case, in sui-indexer-alt/sui-graphql-rpc/sui-package-resolver) which helps to orient quickly (the reason I noticed this was that I looked for ProcessedObjInfo at the top of the file when I saw it mentioned and I couldn't find it there).

It's a forcing function for splitting up files, when they get large enough that a type is far away from its impl block -- otherwise it becomes very easy to concatenate together multiple clusters of types with their impl blocks in a file that grows arbitrarily long without noticing, because you view each cluster as its own unit. It only becomes a problem when someone tries to take a holistic view of the file, and that person struggles.

wlmyng · 2024-12-11T01:26:07Z

Am I understanding correctly that this pruning implementation is unique for the obj_info table, and is needed because we cannot directly prune this table based on checkpoint sequence number, as objects in the pruning range may still be considered live object info?

lxfind · 2024-12-11T03:14:43Z

Am I understanding correctly that this pruning implementation is unique for the obj_info table, and is needed because we cannot directly prune this table based on checkpoint sequence number, as objects in the pruning range may still be considered live object info?

Yes that is correct!

lxfind temporarily deployed to sui-typescript-aws-kms-test-env December 6, 2024 22:23 — with GitHub Actions Inactive

vercel bot deployed to Preview – sui-docs December 6, 2024 22:24 View deployment

lxfind requested review from bmwill, amnn, gegaowp and emmazzz December 6, 2024 22:38

lxfind force-pushed the indexer-alt-add-obj-info-pruner-pipeline branch from 37e5b35 to 3af1bd2 Compare December 8, 2024 21:59

lxfind temporarily deployed to sui-typescript-aws-kms-test-env December 8, 2024 21:59 — with GitHub Actions Inactive

vercel bot deployed to Preview – sui-docs December 8, 2024 22:00 View deployment

lxfind force-pushed the indexer-alt-add-obj-info-pruner-pipeline branch from 3af1bd2 to 6ecab30 Compare December 8, 2024 22:30

lxfind temporarily deployed to sui-typescript-aws-kms-test-env December 8, 2024 22:30 — with GitHub Actions Inactive

vercel bot deployed to Preview – sui-docs December 8, 2024 22:31 View deployment

amnn approved these changes Dec 9, 2024

View reviewed changes

lxfind force-pushed the indexer-alt-add-obj-info-pruner-pipeline branch from 6ecab30 to 929071d Compare December 10, 2024 01:19

lxfind temporarily deployed to sui-typescript-aws-kms-test-env December 10, 2024 01:19 — with GitHub Actions Inactive

vercel bot deployed to Preview – sui-docs December 10, 2024 01:20 View deployment

lxfind force-pushed the indexer-alt-add-obj-info-pruner-pipeline branch from 929071d to 167f008 Compare December 10, 2024 01:21

lxfind temporarily deployed to sui-typescript-aws-kms-test-env December 10, 2024 01:21 — with GitHub Actions Inactive

lxfind enabled auto-merge (squash) December 10, 2024 01:22

vercel bot deployed to Preview – sui-docs December 10, 2024 01:24 View deployment

lxfind disabled auto-merge December 10, 2024 01:46

[indexer-alt] Add obj_info_pruner

67dcb2a

lxfind force-pushed the indexer-alt-add-obj-info-pruner-pipeline branch from 167f008 to 67dcb2a Compare December 10, 2024 01:48

lxfind enabled auto-merge (squash) December 10, 2024 01:48

lxfind temporarily deployed to sui-typescript-aws-kms-test-env December 10, 2024 01:48 — with GitHub Actions Inactive

vercel bot deployed to Preview – sui-docs December 10, 2024 01:50 View deployment

lxfind merged commit 3265cb7 into main Dec 10, 2024
50 of 52 checks passed

lxfind deleted the indexer-alt-add-obj-info-pruner-pipeline branch December 10, 2024 02:24

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

[indexer-alt] Add pruner pipeline for obj_info #20539

[indexer-alt] Add pruner pipeline for obj_info #20539

lxfind commented Dec 6, 2024 •

edited

Loading

vercel bot commented Dec 6, 2024 •

edited

Loading

amnn left a comment

amnn Dec 9, 2024

lxfind Dec 9, 2024

amnn Dec 9, 2024

lxfind Dec 9, 2024

lxfind Dec 10, 2024

amnn Dec 9, 2024

amnn Dec 9, 2024

amnn Dec 9, 2024

lxfind Dec 9, 2024

amnn Dec 10, 2024

wlmyng commented Dec 11, 2024

lxfind commented Dec 11, 2024

		// TODO: We could consider make this more efficient by doing some grouping in the collector
		// so that we could merge as many objects as possible across checkpoints.

[indexer-alt] Add pruner pipeline for obj_info #20539

[indexer-alt] Add pruner pipeline for obj_info #20539

Conversation

lxfind commented Dec 6, 2024 • edited Loading

Description

Test plan

Release notes

vercel bot commented Dec 6, 2024 • edited Loading

amnn left a comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

wlmyng commented Dec 11, 2024

lxfind commented Dec 11, 2024

lxfind commented Dec 6, 2024 •

edited

Loading

vercel bot commented Dec 6, 2024 •

edited

Loading