This html link extractor reads pasted HTML and lists each anchor that has an href. The columns are the href attribute, the anchor text, rel, and target. The parse is DOMParser with text/html. The nodes stay in that document. They are not inserted into this page, and the paste is not assigned to innerHTML.
The href is the attribute as written. A relative path, a protocol-relative URL, a fragment, a mailto address, and a javascript: URL all stay in that form. The page does not read the anchor’s resolved URL property, because that property would prefix this site’s address onto a relative path. It also ignores a base element in the paste. Nothing is resolved against the address bar.
Anchor text is the text content of the link. Nested tags contribute their text, so a strong element inside the anchor is not shown as markup. An image-only anchor has empty text. The image alt is not treated as the label. rel and target are included only when those attributes exist. A missing one is an empty cell, not a guessed value.
Duplicate hrefs stay on separate rows, in document order. The page does not unique them. Anchors that appear only as characters inside script or style are not anchors, and they are not listed. The HTML parser repairs broken markup before the list is built. This is the tree the browser parsed. It is not a token-by-token reading of the source, and it does not promise that attribute order matches the file.
Empty input is an error. HTML with no a that has an href is an error, and any previous rows are cleared. More than 2,000 anchors is an error. The page does not keep the first 2,000 and drop the rest.
<a href="/about">About</a> produces href /about and text About. https://example.com, ../docs/page.html, #section, //cdn.example.com/file, and mailto:test@example.com stay spelled that way. javascript:alert(1) is text in the href cell. It is not run.
An anchor written as <a href="/x"><strong>Hello</strong> world</a> has the text Hello world. Tom & Jerry in the source is the text Tom & Jerry, because the parser decodes the entity into text. Two anchors with the same href are two rows. A fake anchor sitting inside a script element is absent.
The tool does not check that a href is a useful address, does not decode a query string, and does not encode a path. Those are other jobs. This page stops at the attribute text and the visible words of the anchor.
Copy uses the clipboard when the browser allows it. If it does not, the status tells you to copy the table yourself. Clear drops the paste and the rows. There is no history and no preview frame.
Extracts a[href] from HTML parsed in this browser. href, rel, and target are the attribute text, not a resolved URL. Relative links are not turned into this site’s address and nothing is opened or fetched. Anchor text is plain text, not markup. The HTML parser repairs malformed markup; this is not a source-token listing. Enter in the box starts a new line. Limit: 200,000 characters and 2,000 anchors.
No. The href cell is the attribute text. A path such as /about stays /about. This page does not use the page address as a base.
No. Nothing is fetched, and the browser is not sent to the href. A javascript: href stays text in the table.
Text inside script or style is not an element. The HTML parser does not turn that text into an anchor, so this page does not list it.
No. Anchor text is the text of the link. An image-only anchor has empty text. The alt is not copied in.