Jump to content
TNG Community

Drop in # of Indexed Pages


sjwinslow

Recommended Posts

When I started my site with a new domain a few months ago I was expecting a gradual increase in the number of pages being included in Google's index. However, after three months the actual number of indexed pages is less than two months ago. When I started Adsense in mid-Feb. the ad targeting was very poor and I opened up the robots.txt file to basically allow Google access to all pages on my site. Within a couple of weeks Google indexed 17,000 pages total and even the getperson.php pages had about 360 entries in Google's index. However, Google was killing my bandwidth with over 100,000 pages per month being crawled.

Toward the end of March I placed the more restrictive robots.txt file back in service. This more restrictive file blocked access to all but three TNG links (surname-all, search & getperson). Now I expected the 17,000 total pages to drop but I expected the getperson pages to continue to increase. NOTE: Even when the robots.txt file was not restricting Google, the sitemap still only included the paths associated with the three pages of TNG (surname-all, search & getperson). The sitemap was updated on 4/9 to include all getperson pages. The prior sitemap only had about 3,000 out of the 10,000 total getperson pages.

Once the restrictive robots.txt file was re-installed the total pages indexed dropped very quickly. In the last week of March total pages went from over 17,000 to less than 2000. The surprising thing was the getperson pages indexed went from 360 to less than 200. Google continues to crawl my getperson pages at a rate of about 100 per hour. It wasn't until the 5th of April that I started logging the total and the getperson pages that were included on Googles index. The log entries are below:

Date Tot Site getperson

4/5 737 120

4/6 560 72

4/7 558 70

4/10 545 97

4/11 699 94

4/12 783 99

4/13 959 114

4/14 890 67

4/17 627 61

4/18 469 61

4/19 464 56

Any guesses as to what is going on?

Steve

Link to comment
Share on other sites

  • Replies 97
  • Created
  • Last Reply

Top Posters In This Topic

  • sjwinslow

    47

  • Ed Barnard

    16

  • arnold

    11

  • theKiwi

    6

What you are experiencing is what is happening to a lot of us. It's how Google does business. If you write to Google and ask why their robots visit so often and index so infrequently, the standard reply is that not every Google robot visit results in the page being indexed. The reply does not go on to explain the huge disparity in number of visits vs. number of pages indexed. This imbalance will continue to be a fact of life until enough of us complain to goggle.

I went to your showlog.php and saw that it lists visit after visit by search engine robots. If you want to allow robots to continue visiting, but not have them listed on your log, please go to this post on the TNG Forum: showlog.php and search engine robots

Link to comment
Share on other sites

Arnol,

Thanks for the tip to not log all the search engine visits in the TNG log. I'm keeping an eye on them now but I'm sure I'll soon tire and want to view more useful users.

I guess what surprised me was the drop in getperson pages indexed. I'm just baffled as to why if a page has been crawled, indexed and still accessible, why would Google remove it from its index. I know there is no definite answer because we won't get one from Google. But guess I'm looking for speculation here as to what would cause these actions by Google.

The monthly bandwidth remaining on my hosting account is fairly low right now. If I put the unrestricted robots.txt file in service, it might be interesting to see what happens. I would speculate that the overall indexed pages should go up but I can't come up with a good reason why the getperson pages would.

Steve

Link to comment
Share on other sites

If you spend any time researching how google works, you soon encounter the concept of PageRank. It has more to do with who is linking to whom rather than on the actual content of pages visited. All those visits from googlebot are looking not just for subject matter to index, but more importantly googlebot is tallying who you are linking to and who is linking to you.

Here's an example going around the internet these days. Type the word "failure" into the google search page and click "I'm feeling lucky". The resulting page does not have the word "failure" anywhere within it -- it comes up because of a legion of bloggers using the word "failure" when describing (and linking to) the subject of that page.

Mind you...google will say that PageRank is only one of several factors influencing search result placement. But they're never going to spell it out explicitly since those are trade secrets - why would they give the competing search engines a leg up?

As for the disparity between googlebot visits and page listings... you can always block the bot from all but your homepage. Why do you need thousands of your pages listed on google search results? (And what search term are you using to find them?) Wouldn't it be better to decide on the key search terms or phrases you want to compete for and then work to improve your ranking on those? Steve, you already have a top ranking on the phrase "winslow genealogy" in google.

Link to comment
Share on other sites

Why do you need thousands of your pages listed on google search results?

Robert,

It is very important to us to have thousands and thousands of our pages listed on Google. It is important to us to have all of the 177,547 deceased individuals in our database listed in Google's index. That is the only way, when someone is looking for a particular person, that a search on Google will reveal that our website has information on that person. Further, of those 177,547 individuals, only 20,515 have the surname of Sprague. Our keywords mention "Sprague," but not the many, many other surnames.

Our website is unusual in that we address the female lines as completely as we do the male lines. With that in mind, it is clear why the surname Sprague is not such a numerically prominent surname in our database.

Link to comment
Share on other sites

rdmorss,

Hmmm.... Well I must admit it will be very difficult to obtain multiple links to the individual getperson pages to raise their page rank. (Although Darrin's GENDEX site should provide at least one :-D ) Speaking of which Darin's genealogy site has more getperson pages indexed (17,000) than he has people (14,300) in his database.

And I agree I can't complain about the placement of my home page (or several of the html pages on my site) for any of the keywords that are associated with my site. But the primary reason my site exists is so other genealogist who need the information contained on it can use it. Maybe it's naive of me but when I am working on some of the hard-to-fine individuals in my tree I will search for their name using Google. There are a lot of web sites that don't submit their gedcoms to the big genealogy firms but post them directly on the web (mine included). In fact about 10% of the hits I get now on my site are directly to getperson pages (it use to be much higher). The primary reason to get all these pages indexed is so other people can find them. That is also the reason I added information to the title that included birth/death dates and locations.

It actually works quite well when the page is indexed. Try Google searching for Lydia Cobb born in Middleboro "Lydia Cobb Middleboro" CLICK HERE

However, this response is a little off topic and based on Darrin's site having all of its pages indexed, I don't believe page rank is the reason mine are being removed. I believe it must be more associated with the restriction I placed on the bots not being allowed to crawl all of my TNG pages. I have never heard of anyone being penalized by reducing the quantity indexed pages because they didn't allow full access to all pages. But at this point in my assessment it looks like a possibility.

Steve

P.S. As a test I have allowed bot full access to my site for the remainder of this month. We'll see if there is any change in the number of getperson pages indexed.

Link to comment
Share on other sites

That is also the reason I added information to the title that included birth/death dates and locations.

Steve,

Would you please share the code for this and which file it applies to? What you have done is a nice addition.

Link to comment
Share on other sites

Steve,

Would you please share the code for this and which file it applies to? What you have done is a nice addition.

Yes, the instructions to do this got buried in the Meta Tag discussion. To find it CLICK HERE

Steve

Link to comment
Share on other sites

I have some weird results as well. On ewbarnard.com, I have 4000+ pages in google - but those are very old pages. Here's what I see at the moment:

The old html pages - they're in google.

The blog at the domain root - it's in google.

The newer html pages (about 3 months old) - they are not in google.

The coppermine photo album - not in google.

The TNG pages - not in google.

I have three sitemaps registered with google:

1. The blog at the domain root ewbarnard.com/

2. The coppermine photo album ewbarnard.com/album/

3. TNG pages ewbarnard.com/tng/

and google has read the sitemaps for (2) and (3) but has no indication those pages have been crawled.

Most of those 4000+ pages which ARE in google, are not in any of my registered sitemaps.

I'm wondering if my (2) and (3) sitemaps are just plain being ignored because of the existence of the (1) domain root sitemap.

Link to comment
Share on other sites

Jodi Sweere

My website was online for a over full year before I felt it was at a level I wanted it to be. I changed it often within that first year, too, since I was trying to work out how to accomodates my needs and make it more intuitive for users.

The site itself was crawled often even without listing it anywhere. I was already placing in the top ten on Yahoo and sat around 50 on Google pagerank-wise.

In January I began to aggressively list the website. I began with the TNGLinks RSS feed and TNG Ring to increase linking to my site. I also registered with Darrin's TNGNetwork to give search access to my data. That same week I generated a sitemap to Google and placed Google Adsense on the homepage and a few of the static html pages (external to TNG) I had linked from the homepage. The Adsense content was immediately relevent to genealogy, but so is my homepage content, so I didn't expect any problem in that area.

Within days googlebot was crawling my site like crazy and other previously unseen bots showed up as well. Within 2 weeks my page rank on a Google search for my main surname went from 50-something to #1. Another thing worth mentioning is that after I added the Adsense my Yahoo rank dropped way back for a month or so. It's now back to #1 on my main surname search. As far as I can tell, none of my getperson pages are being ranked. My robot.txt file allows indexing of them.

In the time since, I've also placed my page with several genealogy related sites, such as Cindy's List. I plan on doing more of that, as time permits.

I should also mention that my site is gatewayed through WordPress, and my page rank drops off (especially on Yahoo) if I haven't made an update over a few weeks, so I've tried to write small blurbs with site related content to keep it current. Also, using a blogging software as the homepage allows me to submit the site to other blog related listing sites, such as Technorati.

I've had cousins from collateral lines find me, although I didn't ask how, and perhaps I should have. I do the same as someone else mentioned...periodically I Googlesearch names for my database, often spending hours diggin deeper into the less high-ranked pages.

I'm wondering if Google holds off listing indexed pages for a period of time after a newly submitted site has been indexed for weeks on end, in order to track changes? Perhaps the subsequent indexing is used for comparative purposes.

Link to comment
Share on other sites

Jodi,

I recently posted on the mail list that I've been reading a lot about the "sandbox effect" that Google has for new sites. This seems to be a common problem where Google may crawls your site like mad for months before they actually display your pages in their index.

The claim is that this process is used to try to cut down on the spam sites that Google shows in their index. Google started this process in early 2004 and most (but not all) sites started after that time get caught in the sandbox effect.

In the articles I've read no one has come up with a workable solution to ensuring an early exit from the sandbox. Also the duration of the probation seems to vary from as little as 3 months to as much as 15 months. During this probation period you may have you home page indexed and a handful of others but most pages can't be found until your site exits the sandbox.

I have noticed that all of my html pages and the pages in my phpbb froum are indexed but relative few of the TNG pages (as I've mentioned in earlier posts). I have the robots.txt file opened up to allow Google full access to my TNG files in the hopes that some of the earlier pages will come back. It's still too early to tell if there is a trend with that effort.

Steve

Link to comment
Share on other sites

It's been a week or so since I started watching the number of my genealogy pages being indexed by Google. During that time the numbers seem to have stabilized but at a much lower level than I had hoped. Below is a chart that plots both the number of total pages indexed (the blue "Site" line) and the number of getperson pages (the red "Inurl" line) during the last coupl of weeks.

IPB Image

The number of individuals reachable using getperson pages is about 10,000 on my site while the number indexed on Google is just over 100.

One of the observations I made earlier was that Google was indexing the most frequently visited pages. However, with more study this speculation did not match the pages being indexed. It especially didn't explain why Google indexed all of my html pages and even my phpbb forum.

I believe now that Google is not displaying pages that appear to it as being similar. If you look at the layout of most of our pages, they have 90% the same information as others we display. Some of the latest thinking (check out Webmaster World) is that sites are currently being reviewed by Google with a new ranking process. This process retains only the pages that have significantly different content.

When we look at our pages, we focus on the specific few areas that has unique information. The search bot looks at the complete page and compares it to other pages and determines if there is enough differences to warrant it being displayed in the index.

I've noticed that the getperson pages on my site that are most likely to be indexed are the ones with notes and sources. This additional information makes these pages unique from the others. This may also explain why the html and the forum on my site are also indexed but most of the TNG pages are not.

However, this explanation falls short of explaining the sites that have no TNG pages in the index or why some of the older sites (like Darrin's) have almost all of their pages indexed. I believe the factors that Google uses in determining some of these other situations are more complex than this simple explanation I have provide here. Some of those include the age of the site and the number of relevant links you have to your site. Since Google is not talking we'll never know for sure.

Steve

Link to comment
Share on other sites

I agree with what you are saying. One note, when tracking the number of indexed pages, I think you might have to watch out which Google server is being 'asked' for how many instances of site there are.

For half a day, I had only 4 pages index, when I had 496 just before :shock:

Then it went back to normal...

I didn't even really consider the similarity 'penalty' before. It would near impossible to guess at what threshold Google consider a page similar, but you might have a point. Though, I have quite a bit of 'full source' text on people. I can't see how Goolge could consider a few words like a surname repeated 5 or times combined with paragraphs of unique text similar to any other page. :-?

Some interesting things:

- I have new dynamic pages (not TNG) that were spidered and indexed in a matter of 2 or 3 days.

- I have very old dynamic pages (non TNG) that have not existed for 7 - 8 months, yet still remain indexed.

- I had a couple getperson.php pages index after being spidered for 4 or 5 months. After changing over to a CMS and the URL changed, they dropped out of the index almost immediately. Ok, not immediately, but in less than a week and they were gone.

I think Google does this to get a good yuk watching us try to figure it out :razz:

Rush

Link to comment
Share on other sites

I can't see how Goolge could consider a few words like a surname repeated 5 or times combined with paragraphs of unique text similar to any other page.

Rush,

When I look at the keywords that Google sees for my site (This is for all of my pages; html, forum & TNG) it see such words as search, location, events & sheet in the top 10 key words for my whole site (this is taken from the page analysis done by sitemap).

These are all the "title" words being used by the TNG template getperson pages. If you look at a displayed getperson page, the common template contains dozens of words that are repeated literally thousands of times by getperson pages displayed by our database. When the bot looks at the page it sees all of these template title words and includes them in the calculations for determining the similarity of pages. In reality there is very little difference in our pages when the total page is looked at.

As I said earlier this doesn't explain all of the cases of not being indexed by Google but I was able to show a very close correlation between the pages indexed on my site and the pages that contained very unique information.

I'm not sure what if anything we can do about this situation if it exists. The latest thing I'm watching is that the keywords that Google displays in its sitemap analysis seems to be continually changing. Some of the words that use to be associated with my site, birth, sex & died (the most frequent getperson template words) no longer are included in the top 10. This may mean that Google's bot is figuring out that these words are not unique content. It may mean that over time the Title words drop in importance and Google only uses the content words to determine if the page should be indexed. This could explain why the older sites have more pages indexed.

It is an interesting hypothesis but it hard to believe that the Google bot is that smart. On the other hand its also hard to explain why birth, sex and death are no longer considered frequent words for my site, there all still there.

Steve

Link to comment
Share on other sites

I think that it is time for Darrin to jump in and address this thread.

A concern of mine when we chose to incorporate TNG into our website was the possibility that our many, many web pages would no longer be picked up by Google and others. Darrin assured me that my concern was without merit.

However, my experience is that Google continues to hit us hundreds to thousands of time a day, while hardly indexing us at all. We made the switch to TNG in August, 2005. Before that date, web pages for individuals in our website were easy to find via Google. That is no longer true.

Below is information found on the TNG website pertaining to search engines.

TNG Features

Still visible to search engines

Even though TNG pages are not created until requested, they can still be indexed by external search engines like Yahoo, Google, Excite, Lycos, etc., etc. That's because the finished pages are still just HTML, and the engines have to request the pages the same way you or I would. Just remember that it sometimes takes a while for the robots & crawlers to find your pages. Don't want them to find you? Just put the appropriate "no index, no follow" meta tag in your custom header.

Link to comment
Share on other sites

Steve,

Regarding the title tag, I create those dynamically, so rarely two titles would ever be similar except for surname. Those, like you said, still might get 'flagged' by Google as duplicate pages. I think one of the problems with the way TNG does meta tags on a vanilla install is the relevancy is really low. Doing dynamic meta tags pushes the relevancy up over 80%.

I suppose the usefulness of meta tags now can be debated. I don't think they are as 'powerful' as they use to be. Backlinks and Page Rank is probably where its at for SEO.

Only Google could give us a definitive answer and their not talking ;)

As a test I thought of using a script to create a batch of html pages from the getperson to see how quickly Google indexes them. If I do them for one surname, Google should be really slow to index as its trying to deal with what it might perceive as duplicate content.

Rush

Link to comment
Share on other sites

Rush,

I guess I shouldn't have used the word "title" in the context that I did. I was really referencing all the words in the getperson pages that prefix our genealogy data ex., Birth 4 Jun 1851 Death 15 Apr 1893 not the meta title tags. There are just so many words that are repeated on our pages over and over again that its no wonder that Google see them as duplicate pages.

If my hypothesis is correct it should not help to convert the php getperson pages to static html pages. The html page will also contain the same repetitive words. It would be and interesting experiment though.

Steve

Link to comment
Share on other sites

...A concern of mine when we chose to incorporate TNG into our website was the possibility that our many, many web pages would no longer be picked up by Google and others. Darrin assured me that my concern was without merit...

arnold,

Sorry I didn't see your post earlier. I understand your frustration as I too feel the same thing. We have put a lot of time and effort in getting our information out on the web with the intent of people being able to locate the ancestors they're searching. But in fairness to Darrin, I honestly believe that the search engines work differently now than they did even six months ago.

It will be very interesting to see the results of Rush's test he is proposing to determine if the same information in html format will be indexed better than it is in php format. If he can convert 100 pages and only 1% are indexed after 1 month then we know that php is really not a factor and it just the duplicate nature of our genealogy pages. However, if 90% are available then I think we have some serious decisions to make if we want to stay with the php engine of TNG.

I'm sure a lot of people are not concerned about being fully indexed by search engines and this test will have no effect on their use of TNG. But where only 100 out of 10,000 of my pages are currently available for searching it is a major concern of mine. I know you have even more pages and a smaller percentage that are indexed and are rightly concerned.

Steve

Link to comment
Share on other sites

I'm seeing the same thing as everyone else. When I initially uploaded the sitemap, google went crazy on my site. Within two weeks I had 18,000 links listed in google (I have ~7000 people in my database). It has now dropped to about 100 links and they are all for inexistant genealogy pages from the software I previously used as well as some pages from my calendar software that I no longer use.

And google continues to spider my site everyday... just nothing to show for it.

And as a side note, even the tng forum site seems to be losing favor with google. At one point it was indexing most of the topics going on and now it doesn't have any of the topics listed.

Link to comment
Share on other sites

Having built several sites for people in the past, whether it was an auction site, collectibles site or a gaming site, the one thing I learned over the years was search engines do not like PHP. You can do things to help your pages be more search engine friendly, but you can only do so much.

Search engines were built with static pages in mind. Companies that build the search engines have changed their engines so it will pickup dynamic sites also, but still static is still king. Also money talks. The top Dynamic sites that you see all the time are sites that bring in income for Google. You wont find a better online software than TNG for genealogy period.

If you really want to get your names into the google index, create your main page in HTML. Somewhere on the page, place a list of the Surnames that you have listed. You don't have to link them, but can if you want. If you have a large number, then just do the top surnames. Once google hits your page, it will spider all your links and find whatever you want it to. The main thing is the index page with you surnames will get listed and whenever someone looks for that surname it will more than likely come up.

Where your site shows on google is determined mostly by link popularity and if your providing income for google. The more people that link to your site, the more it will show up in google and the higher up the ranking it goes. Also you need to factor in the $$$ for google. They may not say it, but it's a business and they like to make money like anyone else.

This is all just from my experience. Your may differ of course. I think a site that provides info should concentrate on quality and not quantity. If it's quality, then more people will visit and spread the word and other sites may link to you.

Just my 2 cents.

Link to comment
Share on other sites

Here's a pretty good tutor on SEO.

If I wanted to improve the search rank, I would probably write more "histories" as static pages and href them close to the top of the index page. Which is probably a good idea regardless of SEOing. The stories and anecdotes gathered during research are a rich part of a family's history and are the easiest things to lose as the storytellers pass on. Census records, etc. will be there regardless of whether the person is alive or not. Not so the stories. Putting them in static pages alongside the php content may take a little more work but might make it worthwhile in not only SEO, but in retaining visitors. People want to be entertained. Give them a little "Pull a chair up to the fire and I'll tell you a story..."

Link to comment
Share on other sites

sjwinslow

Jpaprocki,

All good points and I'm still considering your suggestion to add the surnames to the index page. This may help somewhat even though the surname-all contents of my site are indexed in Google. The problem I'm mulling over is that even for my primary surname (which has a pretty good ranking in Google) if a potential visitor only searches for the surname my listings fall several dozen pages back. Between all the paid ads and business with high ranking on all the surnames I really never hope to compete with them.

My target is more the researcher who knows who they are looking for and will use firstname+surname (or some other combination). From my current stats, I can see that even with so few pages listed on Google, I receive a large percentage of hits from combination names. It would be very difficult to list any significant representation of the individuals on my site on a reasonable sized index page.

OK, back to the Topic of this thread...

Opening up the robots.txt restrictions to allow full access to my site for the last half of April had the reverse desired effect. After holding steady for many days the total number of pages listed for my site dropped from 470 on Friday to 363 today (23% drop). The number of getperson pages even had a more significant percentage drop from 109 to 55 (49% drop). Google had reached a rate of crawling 1920 pages per day before I reset the restrictions. My 5G bandwidth allowance would not have lasted a complete month at that rate. I have placed the restrictions back in the robots.txt in order to see if there will be any recovery of getperson pages. I'm not sure why just allowing Google more access to my site causes it to reduce the number of indexed pages while increasing the number of crawled pages. It almost seems as Google has a "Duplicate Content" value for the whole site and another value for the page. Carrying this assumption further, the site Duplicate value may be used as a threshold for pages being included in the index. If the page has a lot of uniqueness then the page is include in the index. But if more duplicate pages are added to the site then the threshold is increased and only the most unique pages are displayed.

I've also started tracking the top 20 common words the sitemap page analysis provides. I noticed that after this last experiment, several new common words showed up in the top twenty (married, children, name & print) common words on my site. This may just be coincidence and all are found on the getperson page except for "name". The most prevalent place the word "name" is found is on the searchform.php. This is one of the pages previously blocked from Google.

Link to comment
Share on other sites

sjwinslow

Walt,

I agree. It is looking like unique content is the secret here. However, I think I'll take the SEO hit and just keep adding to the notes in the php pages rather than adding html pages. My effort is not just about trying to build the best SEO site and I don't want to do things that are counterproductive to the site organization or people being able to find the material.

I'm hoping the best compromise will be including unique information in the notes, events and sources with each of the individuals on my site. If this gets them in the index, it's that much better. It's just going to take quite a while to research and publish information on 10,000 individuals. I'm hoping that as more individuals on the site are seen as unique then the threshold for the whole site will also come down.

I believe that from now on as new sites are started, their site's inclusion in the Google index is more dependent on the site's content being unique. This may be a hurdle for many genealogy sites (html or php) since the most common form is to present the data in a tabular format that contains a lot of labels that repeat.

It might be an interesting experiment to use another page format other than the getperson page. An older software package I used once presented a very simple page with the primary couple in the middle of the page, with their parents above and children below with only birth and death information for the primary couple displayed. More detailed information for each person could be locate from links associated with each person. The intuitive layout provided almost no extraneous information on the page and to the search bots each page would be almost completely unique. This type of layout might work better as the primary page presented to the bots rather than the getperson page.

Steve

Link to comment
Share on other sites

Ed Barnard

Here are some approaches to consider, then.

1. Change the template around so that the page's unique content (the individual's info) moves to the beginning of the html page. For example, moving the navigation from the left side to the right side, means it comes after the content rather than before.

2. Present a differently-formatted page to the google spider. Leave off the navigation and all fanciness, pretty much leaving only the unique content for that person.

Link to comment
Share on other sites

Join the conversation

You can post now and register later. If you have an account, sign in now to post with your account.

Guest
Reply to this topic...

×   Pasted as rich text.   Paste as plain text instead

  Only 75 emoji are allowed.

×   Your link has been automatically embedded.   Display as a link instead

×   Your previous content has been restored.   Clear editor

×   You cannot paste images directly. Upload or insert images from URL.


×
×
  • Create New...