Docs & Rules

apache/stormGitHubLast refreshed Oct 2, 2026

This is a public, read-only report. Striff reads this repository's docs, turns each sentence that makes a claim about the code into a rule, and checks the rule against the code on the default branch. How this works

5 names in these docs no longer match the code.

master42147bdlisted 13h ago

Is this yours? Install to manage it

Once installed, Striff checks every pull request.

90
82
1 stale
1
1 stale6
1
1
3
17
1
2 stale4
3
1 stale16
4
2
18
5
5
2
2
3
3
3
28 docs, listed Oct 2
apache/stormOpen repository

Names these docs write that the code no longer has5

Each sentence below names a type the default branch doesn't declare, or declares somewhere else. Edit the doc so it matches the code, or bring the type back.

  • Goneorg.apache.storm.daemon.DrpcServer
    [org.apache.storm.daemon.DrpcServer]({{page.git-blob-base}}/storm-webapp/src/jvm/org/apache/storm/daemon/DrpcServer.java):

    org.apache.storm.daemon declares nothing by this name.

    View on GitHub

  • Renamedorg.apache.storm.metrics2.cgroup.CGroupCPU
    org.apache.storm.metrics2.cgroup.CGroupCPU reports metrics similar to org.apache.storm.metrics.sigar.CPUMetric, but for everything within the CGroup.

    Renamed: CGroupCPU is now CGroupCpu, at storm-client/src/jvm/org/apache/storm/metrics2/cgroup/CGroupCpu.java.

    View on GitHub

  • Goneorg.apache.storm.metrics2.filters.StormMetricFilter
    Custom filters can be created by implementing the org.apache.storm.metrics2.filters.StormMetricFilter interface:

    org.apache.storm.metrics2.filters declares nothing by this name. Closest name there: StormMetricsFilter.

    View on GitHub

  • Goneorg.apache.storm.kafka.trident.TridentKafkaUpdater
    You can create an instance of org.apache.storm.kafka.bolt.KafkaBolt and attach it as a component to your topology or if you are using trident you can use org.apache.storm.kafka.trident.TridentState, org.apache.storm.kafka.trident.TridentStateFactory and org.apache.storm.kafka.trident.TridentKafkaUpdater.

    org.apache.storm.kafka.trident declares nothing by this name. Closest names there: TridentKafkaStateUpdater, TridentKafkaState.

    View on GitHub

  • Goneorg.apache.storm.kafka.trident.TridentStateFactory
    You can create an instance of org.apache.storm.kafka.bolt.KafkaBolt and attach it as a component to your topology or if you are using trident you can use org.apache.storm.kafka.trident.TridentState, org.apache.storm.kafka.trident.TridentStateFactory and org.apache.storm.kafka.trident.TridentKafkaUpdater.

    org.apache.storm.kafka.trident declares nothing by this name. Closest names there: TridentKafkaStateFactory, TridentKafkaState.

    View on GitHub

apache/storm — documented rules

90 of 90 rules, printed October 3, 2026.

The sentence in your docs
Read Oct 2
All cluster state writes go through Utils.serialize(...) / Utils.deserialize(...), which in turn delegate to a pluggable SerializationDelegate selected by the storm.meta.serialization.delegate config.
Utils depends on SerializationDelegate
Holds
Read Oct 2
public Timer registerTimer(String name)
TopologyContext has a registerTimer method
Holds
Read Oct 2
public Histogram registerHistogram(String name)
TopologyContext has a registerHistogram method
Holds
Read Oct 2
public Meter registerMeter(String name)
TopologyContext has a registerMeter method
Holds
Read Oct 2
public Counter registerCounter(String name)
TopologyContext has a registerCounter method
Holds
Read Oct 2
public Gauge registerGauge(String name, Gauge gauge)
TopologyContext has a registerGauge method
Holds
Read Oct 2
Custom metrics reporters can be created by implementing org.apache.storm.metrics2.reporters.StormReporter interface or extending org.apache.storm.metrics2.reporters.ScheduledStormReporter class.
StormReporter is a contract
Holds
Read Oct 2
To try and be as fair as possible to users running short lived topologies the FIFOSchedulingPriorityStrategy extends the DefaultSchedulingPriorityStrategy so that any negative score (a.k.a. a topology that fits within a user's guarantees) would remain unchanged, but positive scores are replaced with the up-time of the topology.
FIFOSchedulingPriorityStrategy extends DefaultSchedulingPriorityStrategy
HoldsPR #9093 · Sep 30
Read Oct 2
All autocredential classes that desire to implement the IMetricsRegistrant interface can register metrics automatically for each topology. The AutoTGT class currently implements this interface and adds a metric named TGT-TimeToExpiryMsecs showing the remaining time until the TGT needs to be renewed.
AutoTGT implements IMetricsRegistrant
Holds
Read Oct 2
The [Config](https://javadoc.io/doc/org.apache.storm/storm-client/3.0.0/org/apache/storm/Config.html) class has a method called registerSerialization that takes in a registration to add to the config.
Config has a registerSerialization method
HoldsPR #9093 · Sep 30
Read Oct 2
There's an advanced config called Config.TOPOLOGY_SKIP_MISSING_KRYO_REGISTRATIONS.
Config has TOPOLOGY_SKIP_MISSING_KRYO_REGISTRATIONS
HoldsPR #9093 · Sep 30
Read Oct 2
When the fallback is enabled, the bridge can be constrained with Config.TOPOLOGY_FALL_BACK_ON_JAVA_SERIALIZATION_FILTER, a [JEP-290](https://openjdk.org/jeps/290) serial-filter pattern (e.g. !org.apache.commons.collections4.functors.*;maxbytes=10485760) applied to every ObjectInputStream the bridge uses for deserialization.
Config has TOPOLOGY_FALL_BACK_ON_JAVA_SERIALIZATION_FILTER
Holds
Read Oct 2
Both the HDFS bolt and Trident State implementation allow you to register any number of RotationActions.
HdfsBolt depends on RotationAction
Holds
Read Oct 2
To partition your your data, write a class that implements the ``Partitioner`` interface and pass it to the withPartitioner() method of your bolt.
HdfsBolt has a withPartitioner method
Holds
Read Oct 2
The SequenceFileBolt requires that you provide a org.apache.storm.hdfs.bolt.format.SequenceFormat that maps tuples to key/value pairs:
SequenceFileBolt depends on SequenceFormat
Holds
Read Oct 2
AvroUtils.addAvroKryoSerializations(conf);
AvroUtils has an addAvroKryoSerializations method
Holds
Read Oct 2
Method .setReaderType(), Alternative config name (deprecated) ~~hdfsspout.reader.type~~, Description Determines which file reader to use Set to 'seq' for reading sequence files or 'text' for text files Set to a fully qualified class name if using a custom file reader class (that implements interface org.apache.storm.hdfs.spout.FileReader)
HdfsSpout has a setReaderType method
Holds
Read Oct 2
Method .withOutputFields(), Description Sets the names for the output fields for the spout The number of fields depends upon the reader being used For convenience, built-in reader types expose a static member called defaultFields that can be used for setting this
HdfsSpout has a withOutputFields method
Holds
Read Oct 2
Method .setHdfsUri(), Alternative config name (deprecated) ~~hdfsspout.hdfs~~, Description HDFS URI for the hdfs Name node Example hdfs://namenodehost:8020
HdfsSpout has a setHdfsUri method
Holds
Read Oct 2
Method .setSourceDir(), Alternative config name (deprecated) ~~hdfsspout.source.dir~~, Description HDFS directory from where to read files E.g /data/inputdir
HdfsSpout has a setSourceDir method
Holds
Read Oct 2
Method .setArchiveDir(), Alternative config name (deprecated) ~~hdfsspout.archive.dir~~, Description After a file is processed completely it will be moved to this HDFS directory If this directory does not exist it will be created E.g /data/done
HdfsSpout has a setArchiveDir method
Holds
Read Oct 2
Method .setBadFilesDir(), Alternative config name (deprecated) ~~hdfsspout.badfiles.dir~~, Description if there is an error parsing a file's contents, the file is moved to this location If this directory does not exist it will be created E.g /data/badfiles
HdfsSpout has a setBadFilesDir method
Holds
Read Oct 2
Method .setCommitFrequencyCount(), Alternative config name (deprecated) ~~hdfsspout.commit.count~~, Default 20000, Description Record progress in the lock file after these many records are processed If set to 0, this criterion will not be used
HdfsSpout has a setCommitFrequencyCount method
Holds
Read Oct 2
Method .setCommitFrequencySec(), Alternative config name (deprecated) ~~hdfsspout.commit.sec~~, Default 10, Description Record progress in the lock file after these many seconds have elapsed Must be greater than 0
HdfsSpout has a setCommitFrequencySec method
Holds
Read Oct 2
Method .setMaxOutstanding(), Alternative config name (deprecated) ~~hdfsspout.max.outstanding~~, Default 10000, Description Limits the number of unACKed tuples by pausing tuple generation (if ACKers are used in the topology)
HdfsSpout has a setMaxOutstanding method
Holds
Read Oct 2
Method .setLockTimeoutSec(), Alternative config name (deprecated) ~~hdfsspout.lock.timeout.sec~~, Default 5 minutes, Description Duration of inactivity after which a lock file is considered to be abandoned and ready for another spout to take ownership
HdfsSpout has a setLockTimeoutSec method
Holds
Read Oct 2
Method .setClocksInSync(), Alternative config name (deprecated) ~~hdfsspout.clocks.insync~~, Default true, Description Indicates whether clocks on the storm machines are in sync (using services like NTP) Used for detecting stale locks
HdfsSpout has a setClocksInSync method
Holds
Read Oct 2
Method .withConfigKey(), Description Optional setting Overrides the default key name ('hdfs.config', see below) used for specifying HDFS client configs
HdfsSpout has a withConfigKey method
Holds
Read Oct 2
Method .withOutputStream(), Description Name of output stream If set, the tuples will be emited to the specified stream Else tuples will be emited to the default output stream
HdfsSpout has a withOutputStream method
Holds
Read Oct 2
IcebergWriterBolt writes no WAL entries, because it makes nothing visible.
IcebergWriterBolt may not depend on CommitWal
Holds
Read Oct 2
You can provide all the producer properties in your Storm topology by calling KafkaBolt.withProducerProperties() and TridentKafkaStateFactory.withProducerProperties().
KafkaBolt has a withProducerProperties method
Holds
Read Oct 2
You can provide all the producer properties in your Storm topology by calling KafkaBolt.withProducerProperties() and TridentKafkaStateFactory.withProducerProperties().
TridentKafkaStateFactory has a withProducerProperties method
Holds
Read Oct 2
This class uses a Builder pattern and can be started either by calling one of the Builders constructors or by calling the static method builder in the KafkaSpoutConfig class.
KafkaSpoutConfig has a builder method
Holds
Read Oct 2
This provides a method routedTo that will say which specific stream the tuple should go to.
KafkaTuple has a routedTo method
Holds
Read Oct 2
These interfaces are combined with RedisLookupMapper and RedisStoreMapper and RedisFilterMapper which fit RedisLookupBolt and RedisStoreBolt, and RedisFilterBolt respectively.
RedisLookupBolt depends on RedisLookupMapper
Holds
Read Oct 2
These interfaces are combined with RedisLookupMapper and RedisStoreMapper and RedisFilterMapper which fit RedisLookupBolt and RedisStoreBolt, and RedisFilterBolt respectively.
RedisStoreBolt depends on RedisStoreMapper
Holds
Read Oct 2
These interfaces are combined with RedisLookupMapper and RedisStoreMapper and RedisFilterMapper which fit RedisLookupBolt and RedisStoreBolt, and RedisFilterBolt respectively.
RedisFilterBolt depends on RedisFilterMapper
Holds
Read Oct 2
One subtle aspect of the interfaces is the difference between IBolt and ISpout vs. IRichBolt and IRichSpout. The main difference between them is the addition of the declareOutputFields method in the "Rich" versions of the interfaces.
IRichBolt has a declareOutputFields method
Holds
Read Oct 2
One subtle aspect of the interfaces is the difference between IBolt and ISpout vs. IRichBolt and IRichSpout. The main difference between them is the addition of the declareOutputFields method in the "Rich" versions of the interfaces.
IRichSpout has a declareOutputFields method
Holds
Read Oct 2
[org.apache.storm.spout]({{page.git-tree-base}}/storm-client/src/jvm/org/apache/storm/spout): Definition of spout and associated interfaces (like the SpoutOutputCollector).
SpoutOutputCollector is in org.apache.storm.spout
Holds
Read Oct 2
[org.apache.storm.spout]({{page.git-tree-base}}/storm-client/src/jvm/org/apache/storm/spout): Definition of spout and associated interfaces (like the SpoutOutputCollector). Also contains ShellSpout which implements the protocol for defining spouts in non-JVM languages.
ShellSpout is in org.apache.storm.spout
Holds
Read Oct 2
[org.apache.storm.task]({{page.git-tree-base}}/storm-client/src/jvm/org/apache/storm/task): Definition of bolt and associated interfaces (like OutputCollector).
OutputCollector is in org.apache.storm.task
HoldsPR #9093 · Sep 30
Read Oct 2
[org.apache.storm.task]({{page.git-tree-base}}/storm-client/src/jvm/org/apache/storm/task): Definition of bolt and associated interfaces (like OutputCollector). Also contains ShellBolt which implements the protocol for defining bolts in non-JVM languages.
ShellBolt is in org.apache.storm.task
HoldsPR #9093 · Sep 30
Read Oct 2
Finally, TopologyContext is defined here as well, which is provided to spouts and bolts so they can get data about the topology and its execution at runtime.
TopologyContext is in org.apache.storm.task
HoldsPR #9093 · Sep 30
Read Oct 2
[org.apache.storm.daemon.Acker]({{page.git-blob-base}}/storm-client/src/jvm/org/apache/storm/daemon/Acker.java): Implementation of the "acker" bolt, which is a key part of how Storm guarantees data processing.
Acker is in org.apache.storm.daemon
Holds
Read Oct 2
[org.apache.storm.LocalCluster]({{page.git-blob-base}}/storm-server/src/main/java/org/apache/storm/LocalCluster.java): Utility to boot up Storm inside an existing Java process.
LocalCluster is in org.apache.storm
Holds
Read Oct 2
[org.apache.storm.Thrift]({{page.git-blob-base}}/storm-client/src/jvm/org/apache/storm/Thrift.java): Wrappers around the generated Thrift API to make working with Thrift structures more pleasant.
Thrift is in org.apache.storm
HoldsPR #9093 · Sep 30
Read Oct 2
[org.apache.storm.StormTimer]({{page.git-blob-base}}/storm-client/src/jvm/org/apache/storm/StormTimer.java): Implementation of a background timer to execute functions in the future or on a recurring interval.
StormTimer is in org.apache.storm
Holds
Read Oct 2
[org.apache.storm.daemon.nimbus]({{page.git-blob-base}}/storm-server/src/jvm/org/apache/storm/daemon/nimbus/Nimbus.java): Implementation of Nimbus.
Nimbus is in org.apache.storm.daemon.nimbus
Holds
Read Oct 2
[org.apache.storm.daemon.supervisor]({{page.git-blob-base}}/storm-server/src/jvm/org/apache/storm/daemon/supervisor/Supervisor.java): Implementation of Supervisor.
Supervisor is in org.apache.storm.daemon.supervisor
Holds
Read Oct 2
[org.apache.storm.daemon.task]({{page.git-blob-base}}/storm-client/src/jvm/org/apache/storm/daemon/Task.java): Implementation of an individual task for a spout or bolt.
Task is in org.apache.storm.daemon
Holds
Read Oct 2
[org.apache.storm.daemon.worker]({{page.git-blob-base}}/storm-client/src/jvm/org/apache/storm/daemon/worker/Worker.java): Implementation of a worker process (which will contain many tasks within).
Worker is in org.apache.storm.daemon.worker
HoldsPR #9093 · Sep 30
Read Oct 2
[org.apache.storm.Testing]({{page.git-blob-base}}/storm-server/src/main/java/org/apache/storm/Testing.java): Various utilities for working with local clusters during tests, e.g. completeTopology for running a fixed set of tuples through a topology for capturing the output, tracker topologies for having fine grained control over detecting when a cluster is "idle", and other utilities.
Testing is in org.apache.storm
HoldsPR #9093 · Sep 30
Read Oct 2
The prepare method parameterizes this batch bolt with the Storm config, the topology context, an output collector, and the id for this batch of tuples.
BaseBatchBolt has a prepare method
HoldsPR #9093 · Sep 30
Read Oct 2
The execute method is called for every tuple in the batch.
BaseBatchBolt has an execute method
HoldsPR #9093 · Sep 30
Read Oct 2
Finally, finishBatch is called when the task has received all tuples intended for it for this particular batch.
BaseBatchBolt has finishBatch
HoldsPR #9093 · Sep 30
Read Oct 2
When using regular bolts, you can call the fail method on OutputCollector to fail the tuple trees of which that tuple is a member.
OutputCollector has a fail method
Holds
Read Oct 2
There are three different interfaces for defining aggregators: CombinerAggregator, ReducerAggregator, and Aggregator.
Aggregator is a contract
Holds
Read Oct 2
You can do that with the TridentTopology#merge method, like so:
TridentTopology has a merge method
Holds
Read Oct 2
Storm has [an implementation of a transactional spout]({{page.git-tree-base}}/external/storm-kafka-client/src/main/java/org/apache/storm/kafka/spout/trident/KafkaTridentSpoutTransactional.java) for Kafka.
KafkaTridentSpoutTransactional is a concrete implementation
Holds
Read Oct 2
[KafkaTridentSpoutOpaque]({{page.git-tree-base}}/external/storm-kafka-client/src/main/java/org/apache/storm/kafka/spout/trident/KafkaTridentSpoutOpaque.java) is a spout that has this property and is fault-tolerant to losing Kafka nodes.
KafkaTridentSpoutOpaque is a concrete implementation
Holds
Read Oct 2
The base State interface just has two methods:
State is a contract
Holds
Read Oct 2
Trident provides the QueryFunction interface for writing Trident operations that query a source of state, and the StateUpdater interface for writing Trident operations that update a source of state.
QueryFunction is a contract
Holds
Read Oct 2
You can then get access to the new values stream for further processing via the TridentState#newValuesStream method.
TridentState has a newValuesStream method
HoldsPR #9093 · Sep 30
Read Oct 2
In this case, since this is a grouped stream, Trident expects the state you provide to implement the "MapState" interface.
MapState is a contract
Holds
Read Oct 2
When you do aggregations on non-grouped streams (a global aggregation), Trident expects your State object to implement the "Snapshottable" interface:
Snapshottable is a contract
Holds
Read Oct 2
IBackingMap looks like this:
IBackingMap is a contract
Holds
Read Oct 2
The OpaqueMap, TransactionalMap, and NonTransactionalMap classes implement all the logic for doing the respective fault-tolerance logic.
OpaqueMap is a concrete implementation
Holds
Read Oct 2
The OpaqueMap, TransactionalMap, and NonTransactionalMap classes implement all the logic for doing the respective fault-tolerance logic.
TransactionalMap is a concrete implementation
Holds
Read Oct 2
The OpaqueMap, TransactionalMap, and NonTransactionalMap classes implement all the logic for doing the respective fault-tolerance logic.
NonTransactionalMap is a concrete implementation
Holds
Read Oct 2
The OpaqueMap, TransactionalMap, and NonTransactionalMap classes implement all the logic for doing the respective fault-tolerance logic. You simply provide these classes with an IBackingMap implementation that knows how to do multiGets and multiPuts of the respective key/values.
OpaqueMap depends on IBackingMap
Holds
Read Oct 2
The OpaqueMap, TransactionalMap, and NonTransactionalMap classes implement all the logic for doing the respective fault-tolerance logic. You simply provide these classes with an IBackingMap implementation that knows how to do multiGets and multiPuts of the respective key/values.
TransactionalMap depends on IBackingMap
Holds
Read Oct 2
The OpaqueMap, TransactionalMap, and NonTransactionalMap classes implement all the logic for doing the respective fault-tolerance logic. You simply provide these classes with an IBackingMap implementation that knows how to do multiGets and multiPuts of the respective key/values.
NonTransactionalMap depends on IBackingMap
Holds
Read Oct 2
Trident also provides the [CachedMap]({{page.git-blob-base}}/storm-client/src/jvm/org/apache/storm/trident/state/map/CachedMap.java) class to do automatic LRU caching of map key/vals.
CachedMap is a concrete implementation
Holds
Read Oct 2
Finally, Trident provides the [SnapshottableMap]({{page.git-blob-base}}/storm-client/src/jvm/org/apache/storm/trident/state/map/SnapshottableMap.java) class that turns a MapState into a Snapshottable object, by storing global aggregations into a fixed key.
SnapshottableMap is a concrete implementation
Holds
Read Oct 2
Finally, Trident provides the [SnapshottableMap]({{page.git-blob-base}}/storm-client/src/jvm/org/apache/storm/trident/state/map/SnapshottableMap.java) class that turns a MapState into a Snapshottable object, by storing global aggregations into a fixed key.
SnapshottableMap depends on MapState
Holds
Read Oct 2
Finally, Trident provides the [SnapshottableMap]({{page.git-blob-base}}/storm-client/src/jvm/org/apache/storm/trident/state/map/SnapshottableMap.java) class that turns a MapState into a Snapshottable object, by storing global aggregations into a fixed key.
SnapshottableMap depends on Snapshottable
Holds
Read Oct 2
The bolt interface IWindowedBolt is implemented by bolts that needs windowing support.
IWindowedBolt is a contract
Holds
Read Oct 2
In this case late tuples are going to be emitted on the specified stream and accessible via the field WindowedBoltExecutor.LATE_TUPLE_FIELD.
WindowedBoltExecutor has a LATE_TUPLE_FIELD field
Holds
Read Oct 2
In this case late tuples are going to be emitted on the specified stream and accessible via the field WindowedBoltExecutor.LATE_TUPLE_FIELD. The value of this field is a org.apache.storm.tuple.DetachedTuple, a serializable copy of the original tuple (source component, task, stream, fields and values) detached from the topology context, so that the late tuple stream can also be consumed by bolts running in different workers.
WindowedBoltExecutor depends on DetachedTuple
Holds
Read Oct 2
**Note:** In case of persistent windowed bolts, use TupleWindow.getIter to retrieve an iterator over the events in the window.
TupleWindow has getIter
Holds
Read Oct 2
If the number of tuples in windows is huge, invoking TupleWindow.get would try to load all the tuples into memory and may throw an OOM exception.
TupleWindow has get
Holds
Read Oct 2
The getPartitionPath() method returns a partition path for a given tuple.
Partitioner has a getPartitionPath method
Holds
Read Oct 2
Method .withOutputFields(), Description Sets the names for the output fields for the spout The number of fields depends upon the reader being used For convenience, built-in reader types expose a static member called defaultFields that can be used for setting this
TextFileReader has a defaultFields field
Holds
Read Oct 2
These interfaces are combined with ``RedisLookupMapper` and `RedisStoreMapper` and `RedisFilterMapper` which fit `RedisLookupBolt` and `RedisStoreBolt`, and `RedisFilterBolt`` respectively.
RedisLookupBolt depends on RedisLookupMapper
Holds
Read Oct 2
These interfaces are combined with ``RedisLookupMapper` and `RedisStoreMapper` and `RedisFilterMapper` which fit `RedisLookupBolt` and `RedisStoreBolt`, and `RedisFilterBolt`` respectively.
RedisStoreBolt depends on RedisStoreMapper
Holds
Read Oct 2
These interfaces are combined with ``RedisLookupMapper` and `RedisStoreMapper` and `RedisFilterMapper` which fit `RedisLookupBolt` and `RedisStoreBolt`, and `RedisFilterBolt`` respectively.
RedisFilterBolt depends on RedisFilterMapper
Holds
Read Oct 2
In order to submit a topology as some other user, you can use the StormSubmitter.submitTopologyAs API.
StormSubmitter has submitTopologyAs
Holds
Read Oct 2
Alternatively you can use NimbusClient.getConfiguredClientAs to get a nimbus client as some other user and perform any nimbus action (i.e., kill/rebalance/activate/deactivate) using this client.
NimbusClient has getConfiguredClientAs
Holds
Read Oct 2
Nimbus will invoke the populateCredentials method of all the configured implementation as part of topology submission.
INimbusCredentialPlugin has a populateCredentials method
Holds
90 rules from 17 docs

See these where the change is. The browser extension puts a pull request's doc findings, and a diagram of what it changed, on the GitHub page itself, so the code and what your docs say about it are side by side.