On-Chain Analytics Explained: Glassnode, Nansen & Arkham

Bartek Hagan

(4 days ago)

20 min read

Share:

This guide covers what on-chain analytics platforms actually see, how address clustering turns raw addresses into named entities, and where those labels break. Tested on 10 September 2026, all three commercial APIs refused every request without a key.

On-Chain Analytics Explained: Glassnode, Nansen & Arkham

Introduction

On 10 September 2026 the largest known bitcoin address held 248,597.59 BTC, worth about $19.2bn. Two free services returned that balance to the satoshi in under a second, with all 5,583 transactions behind it. Neither returned a single field naming the owner. That gap is the entire on-chain analytics industry. It is also why a query to Glassnode, Nansen and Arkham on the same day came back HTTP 401, HTTP 401 and HTTP 400, with no data attached. This guide covers what those platforms actually see, how clustering turns anonymous addresses into company names, where that inference breaks, and which questions genuinely require paying for it. The metrics themselves are a separate subject, covered in a companion article.

Key Takeaways

  • Queried without credentials on 10 September 2026, Glassnode returned HTTP 401, Nansen HTTP 401 and Arkham HTTP 400, with no data from any of them.
  • Free public sources returned complete chain history over the same period, including 931,085 market-cap points back to January 2009.
  • The largest known bitcoin address held 248,597.59 BTC, about 1.24% of all mined supply, and no free source returned any owner or label field for it.
  • Entity labels rest on the common-input-ownership heuristic plus off-chain anchors, so a name inherits the uncertainty of every clustering step behind it.
  • Exchange omnibus wallets pool thousands of customers, so a labelled outflow cannot distinguish a withdrawal from an internal rebalance.

What Does the Blockchain Already Publish for Free?

A public blockchain publishes every transaction to anyone who asks. Several services expose that data with no account and no key. Knowing where that free surface ends is the first step in deciding whether a paid platform is worth anything to you. The boundary is sharper than most write-ups suggest.

What the chain publishes for free

Bitcoin's ledger is open, and the services that index it pass most of it straight through. On 10 September 2026, blockchain.com's chart API returned complete daily history with no credentials. That came to 931,085 market-cap points beginning 3 January 2009 and 6,423 daily unique-address readings. Estimated USD transaction value ran 5,835 points back to August 2010 (blockchain.com charts API, 2026-09-10) . mempool.space served twelve consecutive address lookups in a row without a key or a rate limit. Between them that is balances, transaction counts, supply, fees and block data, free and immediate. Nothing in that list is a preview. It is the same ledger the paid platforms read from, delivered in full to an anonymous caller.

Where the free surface stops

The free surface stops at identity, and it stops abruptly. Blockchair's public API returned HTTP 430 after a small number of unauthenticated requests. The response carried an IP blacklist notice directing the caller to apply for a key. That is the ordinary shape of open chain access. Generous on raw data, quick to throttle at volume, and completely silent on who owns anything. The throttling is a capacity problem, solvable with a modest key or a self-hosted node. The silence on ownership is not, because that information was never written to the chain. That distinction runs through the rest of this article.

SourceWhat It AnsweredKey Required
blockchain.com chartsFull daily history for supply, market cap, addresses, feesNo
mempool.spaceAddress balances, transaction counts, mempool stateNo
BlockchairSame class of data, then HTTP 430 and an IP blacklistYes, at any volume
GlassnodeNothing returnedYes
NansenNothing returnedYes

Data current as of September 2026.

Stat cards comparing free chain endpoints that returned full history against three commercial APIs that returned nothing

The commercial platforms sit exactly at that boundary.

What Happens When You Query the Platforms Without a Key?

Three requests, three refusals. Glassnode, Nansen and Arkham each returned an error rather than data when queried without credentials. Set beside the free endpoints above, that is a clear statement of where these companies believe their value sits. It is not in the raw ledger.

Four endpoints, four answers

Glassnode's metrics endpoint returned HTTP 401 with a bare "Authorization Required" page. Nansen returned HTTP 401 with the message "No API key found in request". Arkham returned HTTP 400 and the message "please sign up for an api key" (direct requests, 2026-09-10) . None of the three leaked a sample row, a demo metric or a partial response. The contrast with the free sources is total. One group hands over the raw ledger to anonymous callers without so much as an email address, and the other hands over nothing whatsoever. There is no middle tier in the API surface itself, whatever the marketing pages offer through a browser.

The key wall is the product line

That wall is not an inconvenience to route around. It marks precisely what the platforms consider proprietary. The underlying transactions are public and identical for everyone, so no platform is selling access to the chain. What sits behind the key is the processing: aggregation into named metrics, attribution of addresses to entities, and the historical series that make either usable. Understanding that changes the buying question. It stops being "do I want on-chain data" and becomes "do I want somebody else's interpretation of it". That is far narrower and far easier to answer honestly. Interpretation can be excellent and still be interpretation. The distinction matters most when a number derived from one platform's assumptions gets quoted downstream as though it were a reading off the ledger.

The gap between raw data and interpretation has a concrete size.

Why Can Nobody Free Tell You Who Owns an Address?

The largest known bitcoin address holds nearly $19.2bn. Every free source will tell you its exact balance to the satoshi, its full transaction history and every address it has ever touched. Not one of them will tell you whose it is, and that gap is the entire on-chain analytics business in a single example.

The largest address holds $19bn

Address 34xp4vRoCGJym3xR7yCVPFHoCNxv4Twseo held 248,597.59 BTC across 5,583 transactions, worth about $19.2bn at a price of $77,230 (mempool.space and blockchain.com ticker, 2026-09-10) . That is roughly 1.24% of the 20,082,487 BTC mined to date, sitting at one address, fully visible to anyone. The transparency is complete on every dimension that the ledger records. Balance, history, first receipt, every counterparty address: all public, all verifiable, all free. Anyone can confirm each figure independently in under a minute, from more than one source, and get the same answer.

No free source names it

What no free source returns is a name. mempool.space answers that query with three top-level fields and blockchain.com with seven, and none of them is a label, tag, entity, owner or company. The chain records value moving between keys, not between organisations. Any statement that this address belongs to a particular exchange is an external claim laid over the data. It is exactly the claim on-chain analytics platforms exist to sell. Whether that claim is correct is a completely separate question from whether the balance is correct, and the two get conflated constantly. The balance is a fact about the ledger, checkable by anyone. The owner is an assertion about the world, checkable by almost nobody. Presenting them side by side lends the second the credibility of the first. That is the most common way on-chain commentary misleads without saying anything false.

Stat cards: the largest known bitcoin address holds 248,597.59 BTC worth $19.2bn, and no free source returns a single label field

Turning an address into a name takes a chain of inference.

How Do Platforms Turn Addresses Into Named Entities?

Attribution is built in layers, and every layer is a heuristic rather than a proof. Understanding the sequence is what lets a reader judge how much weight a label deserves. Confidence in a name is inherited from the weakest step in the chain, not the strongest.

The common-input heuristic

The foundation is the common-input-ownership heuristic. When several inputs are spent together in one transaction, whoever signed it must have held the keys to all of them. Those addresses are therefore presumed to belong to one controller. Repeated across millions of transactions, that presumption grows clusters of addresses that move as a unit. A second layer adds change-address heuristics, which guess which output of a transaction returned funds to the sender. Neither step involves any information from outside the chain, and neither is certain. Both are good statistical bets that are wrong some of the time, and nothing in the output distinguishes the confident cases from the marginal ones.

From cluster to name

A cluster is still anonymous. The name comes from off-chain evidence. A deposit address a company published, an address a researcher transacted with deliberately, a court filing, a hack disclosure, a scraped support page. Those anchors attach an identity to one address, and clustering propagates it across everything that address is linked to. The strength of a label depends on two things. The quality of the original anchor, and every clustering step between that anchor and the address in front of you. Platforms rarely show either. A label appears as a clean company name in the interface. Nothing indicates whether it sits one hop from a published deposit address or forty hops from a scraped forum post.

StepWhat It UsesWhat Can Go Wrong
1. Cluster by common inputsInputs spent together in one transactionCoinJoin and collaborative transactions merge unrelated parties
2. Refine with change heuristicsOutput patterns, script types, round amountsWallet software changes behaviour and the guess flips
3. Anchor to an identityA published address, a filing, a disclosureThe anchor may be stale, second-hand or simply wrong
4. Propagate the labelEvery address linked to the anchorOne bad link contaminates an entire cluster

Data current as of September 2026.

Flowchart: a raw address passing through common-input clustering, change heuristics and off-chain evidence to become an entity label that remains an assertion

Some structures defeat the whole chain of inference.

Where Do Entity Labels Get Things Wrong?

Labels fail in predictable places, and the failures cluster around exactly the entities people most want to track. Exchanges, custodians and pooled investment products are the hardest cases rather than the easiest, which is the opposite of the impression a clean dashboard gives.

Omnibus wallets break the map

A large exchange holds customer coins in shared wallets rather than one address per customer. A single labelled address therefore represents thousands of unrelated beneficial owners. A transfer out of it may be a customer withdrawal, an internal rebalance, a move to a new custody provider or a cold-storage rotation. The chain shows an identical transaction in all four cases. Every headline about coins "leaving exchanges" rests on telling those four apart, and the raw data does not support the distinction. The metric is real and the movement happened. The story attached to it is supplied by the person writing the headline.

Labels go stale quietly

Attribution is a snapshot of an arrangement that changes. Exchanges migrate custody providers, funds change administrators, projects rotate treasury wallets, and companies merge or fail without announcing it. A label attached in one year can describe a relationship that ended in the next, and nothing on the chain announces the change. Stale labels are worse than missing ones. A missing label is visibly missing; a stale one reads as current. That is the strongest practical argument for treating any single labelled flow as a lead worth investigating rather than as a finding worth publishing. Repeated flows between the same clusters over time are more trustworthy than any single transfer. A persistent pattern is less likely to rest on one bad link.

The three best-known platforms approach this differently.

What Are Glassnode, Nansen and Arkham Each Built For?

The three tools most often named together are not really competitors doing the same job. They differ in their unit of analysis, and that difference decides which questions each can answer at all, long before any comparison of features or price enters the picture.

Three tools, three shapes

Glassnode is built around aggregate market metrics. Its unit is the network or the cohort rather than the wallet, and its output is a time series describing the whole chain rather than any participant in it. Nansen is built around labelled wallets, primarily on EVM chains. Its unit is the address, with a behavioural tag attached. Arkham is built around entity attribution and the address graph. It aims to resolve activity to named organisations across chains. Those are three genuinely different products that happen to share an input. Comparing them on a single axis produces the same category error as ranking a telescope against a microscope.

What each was built to answer

The practical consequence is that the same question is easy on one and awkward on the others. A cohort-level question about long-term holders fits an aggregate metrics product. A question about what one labelled wallet did this week fits a wallet-first product. A question about which organisation controls a cluster needs an attribution-first product. CoinPaprika's guide to Arkham covers that platform hands-on, and its guide to tracking smart money covers the wallet-following workflow in practice. Neither is restated here. The point of this section is only that picking a tool before framing the question guarantees fighting the tool afterwards.

PlatformPrimary Unit of AnalysisTypical Question
GlassnodeNetwork-wide and cohort aggregatesHow has holder behaviour shifted across the whole chain?
NansenLabelled wallets, EVM-centricWhat did this tagged address do, and who else moved with it?
ArkhamEntities and the address graphWhich organisation stands behind this cluster of addresses?

Data current as of September 2026.

Cost sits on top of all three.

Is a Free Tier Enough for On-Chain Analysis?

For most questions a beginner actually has, free chain data is enough. The paid tiers become necessary at a specific and identifiable point, and that point arrives considerably later than the marketing suggests. Naming it precisely is more useful than any feature comparison.

What the free tiers really give

Browsing a platform's website is not the same as querying it. All three commercial APIs tested here returned nothing without credentials. Any free access therefore runs through a web interface, with its own limits on history, granularity and export. The genuinely free layer is the one described earlier: full raw chain data from public indexers, complete daily history for network aggregates, and no identity at all. That layer answers a great deal. Supply, fees, transaction counts, address activity and every metric derivable from them are reachable without paying anyone, over the full history of the chain.

What paying actually buys

Paying buys three things, and naming them separately helps. It buys entity labels, which cannot be derived from the chain at any price of effort. It buys precomputed metrics with consistent historical series, which saves substantial work without revealing anything new. And it buys coverage across chains that would otherwise mean maintaining several data pipelines. Only the first is genuinely unobtainable elsewhere. The other two are convenience. Convenience is a perfectly reasonable thing to buy, and it earns its price for anyone whose time costs more than a subscription. It is not the same as insight, and a chart nobody else has built is not the same as a fact nobody else can check.

No amount of payment buys the missing half of the picture.

What Can No On-Chain Platform See?

Every platform is limited by what the ledger records, and the ledger records settlement rather than activity. Enormous volumes of economically real trading never touch it at all, which puts a hard ceiling on what any amount of on-chain sophistication can see.

Nothing sees off-chain

Trades on a centralised exchange move balances in that company's internal database. Nothing is written to the chain until a customer deposits or withdraws. The entire order flow of the largest venues is therefore invisible to on-chain analysis. Companion articles on the order book ↗ and on crypto liquidity ↗ cover reading that activity where it actually happens. The same blindness covers over-the-counter desks, internal custodian transfers and layer-two activity that settles in batches. On-chain data is a record of settlement, and settlement is a fraction of what happens. Treating it as a complete view of the market is the largest error in the field.

The exchange ledger problem

This is why exchange-flow metrics are so easily overread. A large transfer into an exchange-labelled cluster is consistent with an intention to sell. It is equally consistent with a custody rotation, a collateral posting or a market maker rebalancing inventory. The chain cannot separate those, and no platform can add information that was never recorded. Intent is not on the ledger. A platform that presents a flow as a directional signal has made an inference. Noticing that the inference is not part of the data is the reader's job. The flow figure may be exactly right and the conclusion drawn from it still unsupported.

Choosing a source starts from the question rather than the tool.

Which On-Chain Question Needs Which Source?

Working backwards from the question eliminates most of the choice before any tool is opened. Only a narrow band of questions genuinely requires a paid platform, and identifying that band is the whole of the decision. Everything else is preference.

Matching question to tool

If the question concerns network-wide activity, the raw chart APIs answer it directly and for nothing. If it concerns the composition of holders or a cohort's behaviour over time, an aggregate metrics product saves real work. The underlying inputs remain public either way. If it concerns a specific named organisation, only an attribution product can attempt it, and the answer arrives as a claim rather than a fact. If it concerns intent, no source answers it. Sorting a question into one of those four boxes takes about a minute and usually removes the need to buy anything at all. The fourth box is the one people skip, and it matters most. A question about motive cannot be answered by better data.

QuestionWhere It Is AnsweredPaid Data Needed
How many addresses transacted today?Public chart APIsNo
What are fees and supply doing over years?Public chart APIsNo
How is holder behaviour shifting by cohort?Aggregate metrics platformConvenience only
Which company controls this cluster?Attribution platform, as a claimYes
Why did they move the coins?NowhereNot obtainable

Data current as of September 2026.

When raw chain data is enough

Most published on-chain commentary rests on metrics computable from free endpoints. A companion article works through those metrics directly ↗, including how they are calculated and where their thresholds stop holding. A reader who wants to check a claim rather than generate one almost never needs a subscription. The exception is any claim that turns on identity. Those are unverifiable without attribution data and, as the previous sections showed, uncertain even with it. That is a narrow exception, and it covers a surprisingly small share of what gets published.

Five-step diagram: state the question, check the raw chain, check aggregate metrics, decide whether identity is required, then pay

That suggests a short test to apply before believing anything.

How Should You Judge an On-Chain Claim?

Three questions separate a checkable on-chain claim from an unfalsifiable one. They apply equally to a platform's own dashboard, a paid research note and an unsourced screenshot on social media. None of the three requires any tooling to ask, and none takes longer than a moment to answer.

A three-question test

Ask first whether the claim depends on a label. If it names a company, it inherits every uncertainty in the attribution chain described above. The person making the claim probably cannot show you the anchor. Ask second whether the underlying number is reproducible from public data. A great many are, and checking takes minutes. Ask third whether the claim asserts intent. "Coins moved to an exchange" is an observation. "Holders are preparing to sell" is an interpretation with no support in the ledger, and the switch between the two usually happens silently inside one sentence.

Reading any on-chain claim

The honest summary is that on-chain analytics is strongest at the aggregate level and weakest at the level people find most compelling. Network-wide counts are solid, reproducible and free. Named-entity flows are the product being sold and the least verifiable thing on offer. Between the two sits a large body of cohort metrics: useful, computable and often oversold. Knowing which layer a claim comes from is more valuable than knowing any individual metric, and establishing it costs nothing. The chain is transparent about value and silent about identity, and almost every on-chain dispute traces back to someone forgetting the second half of that sentence.

Summary

Public blockchains publish everything they record, and free indexers pass it straight through. blockchain.com returned complete daily history without credentials, and mempool.space served twelve consecutive address lookups without a key. Blockchair throttled to an IP blacklist after a few requests, which is a capacity limit rather than a data limit. Against that, all three commercial platforms tested returned errors and nothing else. The wall is not around the chain. It is around the processing applied to it.

That processing is mostly attribution, and attribution is a chain of heuristics. Clustering presumes that inputs spent together share one owner. Change heuristics guess which output returned funds to the sender, and an off-chain anchor supplies the name that then propagates across the cluster. Each step can fail. The failures concentrate on exchanges and custodians, whose omnibus wallets pool thousands of unrelated owners. The three best-known tools differ in their unit of analysis: Glassnode works in network aggregates, Nansen in labelled wallets, Arkham in entities and the address graph. None of them sees off-chain trading, internal exchange ledgers, or intent.

Conclusion

The useful division is between what the ledger records and what someone has asserted about it. Balances, counts, fees and supply are recorded, free, and checkable by anyone in minutes. Names are asserted, sold, and unverifiable from the chain at any level of effort. Most published on-chain commentary rests on the first category while borrowing its authority for the second. A paid platform earns its price when a question genuinely turns on identity, or when precomputed history saves more time than it costs. It earns nothing when the same number already sits in a free endpoint. Sorting which case you are in takes about a minute and is the most valuable habit in this whole field.

Why You Might Be Interested?

If you have read a headline about coins leaving exchanges, it depends on a label that cannot separate a withdrawal from an internal transfer. If you are considering a subscription, check first whether your question is answered by a free endpoint. If you follow named wallets, remember that the name is somebody's inference and the balance is the only part the chain guarantees.

Free sources gave the exact $19.2bn balance of one address and zero fields naming its owner.

Quick Stats

  • 401, 401, 400 — HTTP responses from Glassnode, Nansen and Arkham to unauthenticated requests
  • 931,085 — free market-cap data points returned by blockchain.com with no key, back to January 2009
  • 12 of 12 — consecutive address lookups mempool.space served without credentials or throttling
  • 248,597.59 BTC — balance of the largest known bitcoin address, about $19.2bn at $77,230
  • 1.24% — share of the 20,082,487 BTC mined that sits at that single address
  • 0 — label, tag, entity or owner fields returned for it by any free source checked

Data current as of September 2026.

FAQ

?What is on-chain analytics?

It is the practice of drawing conclusions from data recorded on a public blockchain: balances, transfers, address activity, fees and supply. The raw records are open to everyone. Commercial platforms add processing on top: aggregation into named metrics, and attribution of addresses to real-world entities. The second of those is the part that cannot be reproduced from the chain alone.

?Are Glassnode, Nansen and Arkham free?

Not through their APIs. Queried without credentials on 10 September 2026, Glassnode returned HTTP 401 and Nansen returned HTTP 401 with "No API key found in request". Arkham returned HTTP 400, asking the caller to sign up for a key. Whatever free access exists runs through their web interfaces, with limits on history, granularity and export that vary by product.

?Can you see who owns a bitcoin address?

Not from the chain. The ledger records value moving between keys, never between organisations. Free sources return balances, transaction counts and full histories with no owner field at all. Any name attached to an address comes from clustering plus off-chain evidence. It is an assertion made by whoever attached it, not a property of the data.

?How does address clustering work?

The foundation is the common-input-ownership heuristic. If several inputs are spent together in one transaction, whoever signed it held the keys to all of them, so those addresses probably share a controller. Change-address heuristics refine the picture by guessing which output returned funds to the sender. Both are statistical bets, and both fail on CoinJoin transactions, payment batching and shared custody arrangements.

?Why are exchange flow metrics unreliable?

Because exchanges hold customer coins in shared omnibus wallets. One labelled address can represent thousands of unrelated owners. A transfer out of it is equally consistent with a customer withdrawal, an internal rebalance, a custody migration or a cold-storage rotation. The chain records all four identically, and the interpretation is supplied by whoever writes the headline.

?Do I need a paid on-chain platform?

Only if your question turns on identity, or if precomputed historical series save you more time than the subscription costs. Network-wide questions about supply, fees, transaction counts and address activity are answered completely by free endpoints. Questions about motive are not answered by any source at any price, because intent was never recorded.

?What can on-chain data never show?

Anything that does not settle on the chain. Trading on centralised exchanges moves entries in a private database, so the order flow of the largest venues is invisible. Over-the-counter deals, internal custodian transfers and batched layer-two activity are equally absent. On-chain data is a record of settlement rather than of trading, and settlement is a fraction of total activity.

?How do I check an on-chain claim?

Ask three things. Does it depend on a label, in which case it inherits the uncertainty of the attribution chain? Is the underlying number reproducible from a public endpoint, which many are? And does it assert intent, which no ledger records? A claim that survives all three is worth taking seriously, and most stop at the first.

References / Sources

Sources
  • All endpoint behaviour measured first-hand by direct HTTP request on 10 September 2026. No credentials were used with any service.
  • - blockchain.com: Charts API and Ticker Endpoint (blockchain.com, Sep 2026)
  • - mempool.space: Public Address and Fees API (mempool.space, Sep 2026)
  • - Blockchair: Public API rate-limit response (blockchair.com, Sep 2026)
  • - Glassnode: API authorization response (glassnode.com, Sep 2026)
  • - Nansen: API authorization response (nansen.ai, Sep 2026)
  • - Arkham: API authorization response (arkm.com, Sep 2026)

Related articles

Latest articles

Coinpaprika education

Discover practical guides, definitions, and deep dives to grow your crypto knowledge.

Cryptocurrencies are highly volatile and involve significant risk. You may lose part or all of your investment.

All information on Coinpaprika is provided for informational purposes only and does not constitute financial or investment advice. Always conduct your own research (DYOR) and consult a qualified financial advisor before making investment decisions.

Coinpaprika is not liable for any losses resulting from the use of this information.

Go back to Education