{"id":69602,"date":"2022-07-31T06:09:58","date_gmt":"2022-07-31T04:09:58","guid":{"rendered":"https:\/\/rmarketingdigital.com\/wiki\/robots-txt-2\/"},"modified":"2024-10-28T22:48:28","modified_gmt":"2024-10-28T21:48:28","slug":"robots-txt-2","status":"publish","type":"glossary","link":"https:\/\/rmarketingdigital.com\/en\/wiki\/robots-txt-2\/","title":{"rendered":"Robots.txt"},"content":{"rendered":"<p><\/p>\n<div id=\"mw-content-text\" lang=\"es\" dir=\"ltr\" class=\"mw-content-ltr\">\n<p>The <b>robots.txt file<\/b> is a document that establishes which parts of a domain can be analyzed by search engine crawlers and provides a link to the <i>XML-sitemap<\/i>.\n<\/p>\n<h2>\n<span class=\"mw-headline\" id=\"Estructura\">Structure<\/span><br \/>\n<\/h2>\n<p>The call <i>&quot;Robots Exclusion Standard Protocol&quot;<\/i>, <b>Standard Protocol for Robots Exclusions<\/b>, was first published in 1994. This protocol defines that search engine crawlers must find and read the file named <b>&quot;Robots.txt&quot;<\/b> before starting indexing. That is why it should be placed in the root directory of the domain. In spite of everything, we must remember that not all trackers follow this same rule and therefore, the <b>&quot;Robots.txt&quot;<\/b> they do not promise the 100% access and privacy protection. Some search engines still index blocked pages and even show those with no description in the SERPs. This is particularly the case with websites that contain too many links. Regardless, the major search engines like <b> Google, Yahoo<\/b> and <b>Bing<\/b> yes they conform to the protocol rules <b>&quot;Robots.txt&quot;<\/b>.\n<\/p>\n<h2>\n<span class=\"mw-headline\" id=\"Creaci.C3.B3n_y_control_del_.E2.80.9Crobots.txt.E2.80.9C\">Creation and control of the &quot;robots.txt&quot;<\/span><br \/>\n<\/h2>\n<p>It is simple to create a <b>&quot;Robots.txt&quot;<\/b> with the help of a text editor. At the same time, you can find free tools on the internet that offer detailed information about how to generate a file. <b>&quot;Robots.txt&quot;<\/b> or that they even automatically create it for you. Each file contains 2 blocks. In the first, it is specified for which users the instructions are valid. In the second block the instructions are written, called <i>&quot;Disallow&quot;<\/i>, with the list of pages to be excluded. It is recommended to check carefully that the file has been written correctly before downloading it to the directory since, with basically a tiny syntax error, you can misinterpret the instructions and index pages that, in theory, should not appear in the search results . To check if the file <b>&quot;Robots.txt&quot;<\/b> works correctly you can use Google&#039;s webmaster tool and perform an analysis on \u201estatus\u201c -&gt; \u201eblocked URLs\u201c.\n<\/p>\n<p><a rel=\"nofollow noopener noreferrer\" target=\"_blank\" data-bs-title=\"Archivo:600x400-robotstxt-es-01.png\" data-bs-filetimestamp=\"20180423165522\"><img loading=\"lazy\" decoding=\"async\" alt=\"600x400-robotstxt-en-01.png\" src=\"https:\/\/rmarketingdigital.com\/wp-content\/uploads\/2020\/05\/600x400-robotstxt-es-01.png\" width=\"600\" height=\"400\"><\/a>\n<\/p>\n<h2>\n<span class=\"mw-headline\" id=\"Exclusi.C3.B3n_de_p.C3.A1ginas\">Pages exclusion<\/span><br \/>\n<\/h2>\n<p>The simplest structure of a file <b>robots.txt<\/b> appears as follows:\n<\/p>\n<pre>User-agent: Googlebot Disallow:\n<\/pre>\n<p>This code allows Googlebot to analyze all pages. The opposite, such as the complete ban of the web portal, is written as follows: &#039;\n<\/p>\n<pre>User-agent: Googlebot Disallow:\n<\/pre>\n<p>In the &quot;User-agent&quot; line the user writes who it is addressed to. The following terms can be used:\n<\/p>\n<ul>\n<li> Googlebot (Google search engine)<\/li>\n<li> Googlebot-Image (Google-image search)<\/li>\n<li> Adsbot-Google (Google AdWords)<\/li>\n<li> Slurp (Yahoo)<\/li>\n<li> bingbot (Bing)<\/li>\n<\/ul>\n<p>If the order is addressed to different users, each robot will have its own line. In <i>mindshape.de<\/i> you will be able to find a summary of the most common orders and parameters for creating a <b>robots.txt<\/b>. Furthermore, a link to an XML-Sitemap can be added as follows:\n<\/p>\n<pre>Sitemap: http:\/\/www.domain.de\/sitemap.xm\n<\/pre>\n<h2>\n<span class=\"mw-headline\" id=\"Ejemplo\">Example<\/span><br \/>\n<\/h2>\n<pre>\n# robots.txt for http:\/\/www.example.com\/ User-agent: UniversalRobot \/ 1.0 User-agent: my-robot Disallow: \/ sources \/ dtd \/ User-agent: * Disallow: \/ nonsense \/ Disallow: \/ temp \/ Disallow: \/newsticker.shtml\n<\/pre>\n<h2>\n<span class=\"mw-headline\" id=\"Relevancia_para_el_SEO\">Relevance for SEO<\/span><br \/>\n<\/h2>\n<p>The use of the protocol <b>robots.txt<\/b> influences crawlers&#039; access to the web portal. There are two different commands: <i>&quot;Allow&quot; and &quot;disallow&quot;<\/i>. It is very important to use this protocol correctly since if the webmaster blocks by mistake - by means of the &quot;disallow&quot; command - important files and contents of the web portal, the crawlers will not be able to read or index it. Regardless, if used correctly webmasters are able to inform crawlers on how to review the internal structure of their web portal.\n<\/p>\n<h2>\n<span class=\"mw-headline\" id=\"Enlaces_web\">Web links<\/span><br \/>\n<\/h2>\n<ul>\n<li>\n<a rel=\"nofollow noopener noreferrer\" target=\"_blank\" class=\"external text\" href=\"https:\/\/support.google.com\/webmasters\/answer\/6062608?hl=es\">Information about robots.txt files<\/a> support.google.com<\/li>\n<\/ul>\n<ul>\n<li>\n<a rel=\"nofollow noopener noreferrer\" target=\"_blank\" class=\"external text\" href=\"http:\/\/ignaciosantiago.com\/archivo-robots-txt\/\">Robots.txt file: what it is, what it is for and how to create it<\/a> Blog ignaciosantiago.com<\/li>\n<\/ul>\n<p><!-- \nNewPP limit report\nCached time: 20200502053241\nCache expiry: 86400\nDynamic content: false\nCPU time usage: 0.008 seconds\nReal time usage: 0.009 seconds\nPreprocessor visited node count: 59\/1000000\nPreprocessor generated node count: 108\/1000000\nPost\u2010expand include size: 0\/2097152 bytes\nTemplate argument size: 0\/2097152 bytes\nHighest expansion depth: 2\/40\nExpensive parser function count: 0\/100\n--><!-- \nTransclusion expansion time report (%,ms,calls,template)\n100.00%    0.000      1 - -total\n--><!-- Saved in parser cache with key es_:stable-pcache:idhash:15-0!*!0!!es!5!* and timestamp 20200502053241 and revision id 2769\n -->\n<\/div>","protected":false},"excerpt":{"rendered":"<p>El archivo robots.txt es un documento que establece qu\u00e9 partes de un dominio pueden ser analizadas por los rastreadores de los motores de b\u00fasqueda y proporciona un link al XML-sitemap&#8230;.<\/p>","protected":false},"author":1,"featured_media":69604,"parent":0,"template":"","glossary-cat":[],"class_list":["post-69602","glossary","type-glossary","status-publish","has-post-thumbnail"],"yoast_head":"<!-- This site is optimized with the Yoast SEO Premium plugin v27.0 (Yoast SEO v28.2) - https:\/\/yoast.com\/product\/yoast-seo-premium-wordpress\/ -->\n<title>Robots.txt - R Marketing Digital<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/rmarketingdigital.com\/en\/wiki\/robots-txt-2\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Robots.txt\" \/>\n<meta property=\"og:description\" content=\"El archivo robots.txt es un documento que establece qu\u00e9 partes de un dominio pueden ser analizadas por los rastreadores de los motores de b\u00fasqueda y proporciona un link al XML-sitemap....\" \/>\n<meta property=\"og:url\" content=\"https:\/\/rmarketingdigital.com\/en\/wiki\/robots-txt-2\/\" \/>\n<meta property=\"og:site_name\" content=\"R Marketing Digital\" \/>\n<meta property=\"article:modified_time\" content=\"2024-10-28T21:48:28+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/rmarketingdigital.com\/wp-content\/uploads\/2020\/05\/600x400-robotstxt-es-01.png\" \/>\n\t<meta property=\"og:image:width\" content=\"600\" \/>\n\t<meta property=\"og:image:height\" content=\"400\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/png\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/rmarketingdigital.com\\\/wiki\\\/robots-txt-2\\\/\",\"url\":\"https:\\\/\\\/rmarketingdigital.com\\\/wiki\\\/robots-txt-2\\\/\",\"name\":\"Robots.txt - R Marketing Digital\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/rmarketingdigital.com\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/rmarketingdigital.com\\\/wiki\\\/robots-txt-2\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/rmarketingdigital.com\\\/wiki\\\/robots-txt-2\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/rmarketingdigital.com\\\/wp-content\\\/uploads\\\/2020\\\/05\\\/600x400-robotstxt-es-01.png\",\"datePublished\":\"2022-07-31T04:09:58+00:00\",\"dateModified\":\"2024-10-28T21:48:28+00:00\",\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/rmarketingdigital.com\\\/wiki\\\/robots-txt-2\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/rmarketingdigital.com\\\/wiki\\\/robots-txt-2\\\/#primaryimage\",\"url\":\"https:\\\/\\\/rmarketingdigital.com\\\/wp-content\\\/uploads\\\/2020\\\/05\\\/600x400-robotstxt-es-01.png\",\"contentUrl\":\"https:\\\/\\\/rmarketingdigital.com\\\/wp-content\\\/uploads\\\/2020\\\/05\\\/600x400-robotstxt-es-01.png\",\"width\":600,\"height\":400,\"caption\":\"600x400-robotstxt-es-01-5764978-png\"},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/rmarketingdigital.com\\\/#website\",\"url\":\"https:\\\/\\\/rmarketingdigital.com\\\/\",\"name\":\"R Marketing Digital\",\"description\":\"Agencia SEO | Publicidad redes sociales | Community Management\",\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/rmarketingdigital.com\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"}]}<\/script>\n<!-- \/ Yoast SEO Premium plugin. -->","yoast_head_json":{"title":"Robots.txt - R Marketing Digital","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/rmarketingdigital.com\/en\/wiki\/robots-txt-2\/","og_locale":"en_US","og_type":"article","og_title":"Robots.txt","og_description":"El archivo robots.txt es un documento que establece qu\u00e9 partes de un dominio pueden ser analizadas por los rastreadores de los motores de b\u00fasqueda y proporciona un link al XML-sitemap....","og_url":"https:\/\/rmarketingdigital.com\/en\/wiki\/robots-txt-2\/","og_site_name":"R Marketing Digital","article_modified_time":"2024-10-28T21:48:28+00:00","og_image":[{"width":600,"height":400,"url":"https:\/\/rmarketingdigital.com\/wp-content\/uploads\/2020\/05\/600x400-robotstxt-es-01.png","type":"image\/png"}],"twitter_card":"summary_large_image","schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"WebPage","@id":"https:\/\/rmarketingdigital.com\/wiki\/robots-txt-2\/","url":"https:\/\/rmarketingdigital.com\/wiki\/robots-txt-2\/","name":"Robots.txt - R Marketing Digital","isPartOf":{"@id":"https:\/\/rmarketingdigital.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/rmarketingdigital.com\/wiki\/robots-txt-2\/#primaryimage"},"image":{"@id":"https:\/\/rmarketingdigital.com\/wiki\/robots-txt-2\/#primaryimage"},"thumbnailUrl":"https:\/\/rmarketingdigital.com\/wp-content\/uploads\/2020\/05\/600x400-robotstxt-es-01.png","datePublished":"2022-07-31T04:09:58+00:00","dateModified":"2024-10-28T21:48:28+00:00","inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/rmarketingdigital.com\/wiki\/robots-txt-2\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/rmarketingdigital.com\/wiki\/robots-txt-2\/#primaryimage","url":"https:\/\/rmarketingdigital.com\/wp-content\/uploads\/2020\/05\/600x400-robotstxt-es-01.png","contentUrl":"https:\/\/rmarketingdigital.com\/wp-content\/uploads\/2020\/05\/600x400-robotstxt-es-01.png","width":600,"height":400,"caption":"600x400-robotstxt-es-01-5764978-png"},{"@type":"WebSite","@id":"https:\/\/rmarketingdigital.com\/#website","url":"https:\/\/rmarketingdigital.com\/","name":"R Digital Marketing","description":"SEO Agency | Advertising social networks | Community Management","potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/rmarketingdigital.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"}]}},"_links":{"self":[{"href":"https:\/\/rmarketingdigital.com\/en\/wp-json\/wp\/v2\/glossary\/69602","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/rmarketingdigital.com\/en\/wp-json\/wp\/v2\/glossary"}],"about":[{"href":"https:\/\/rmarketingdigital.com\/en\/wp-json\/wp\/v2\/types\/glossary"}],"author":[{"embeddable":true,"href":"https:\/\/rmarketingdigital.com\/en\/wp-json\/wp\/v2\/users\/1"}],"version-history":[{"count":0,"href":"https:\/\/rmarketingdigital.com\/en\/wp-json\/wp\/v2\/glossary\/69602\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/rmarketingdigital.com\/en\/wp-json\/wp\/v2\/media\/69604"}],"wp:attachment":[{"href":"https:\/\/rmarketingdigital.com\/en\/wp-json\/wp\/v2\/media?parent=69602"}],"wp:term":[{"taxonomy":"glossary-cat","embeddable":true,"href":"https:\/\/rmarketingdigital.com\/en\/wp-json\/wp\/v2\/glossary-cat?post=69602"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}