• 26-01-2007, 17:36:30
    #1
    Misafir
    yeni açtığım program download sitesi devamlı başka sitelerden back alıyo botlar falan indexliyo herhalde bunu nsl engelleyebilirim 20 günlük sitenin 100 küsür backlinki var ve site sandboxtan çıkmıyor.
  • 26-01-2007, 19:41:43
    #2
    Robots.txt yi düzenledinmi arkadaşım ?
  • 26-01-2007, 21:04:24
    #3
    Misafir
    hyr ama Robots.txt e engellermi tam olarak.
  • 26-01-2007, 21:23:28
    #4
    turkdgn adlı üyeden alıntı: mesajı görüntüle
    hyr ama Robots.txt e engellermi tam olarak.
    Robotların ismi belliyse (yani robotlar ismini her seferinde değiştirmiyorsa) robots.txt engeller. Referans olması açısından kendi robots.txt dosyamı veriyorum, aralarında Teleport gibi tüm siteyi indirme programları ve mail toplama programları da var.

    #Banned Robots
    User-agent: psbot
    Disallow: /
    User-Agent: Googlebot-Image
    Disallow: /
    User-agent: CherryPicker
    Disallow: /
    User-agent: EmailCollector
    Disallow: /
    User-agent: EmailSiphon
    Disallow: /
    User-agent: WebBandit
    Disallow: /
    User-agent: EmailWolf
    Disallow: /
    User-agent: ExtractorPro
    Disallow: /
    User-agent: CopyRightCheck
    Disallow: /
    User-agent: Crescent
    Disallow: /
    User-agent: NICErsPRO
    Disallow: /
    User-agent: SiteSnagger
    Disallow: /
    User-agent: ProWebWalker
    Disallow: /
    User-agent: CheeseBot
    Disallow: /
    User-agent: ia_archiver
    Disallow: /
    User-agent: ia_archiver/1.6
    Disallow: /
    User-agent: Alexibot
    Disallow: /
    User-agent: Teleport
    Disallow: /
    User-agent: TeleportPro
    Disallow: /
    User-agent: MIIxpc
    Disallow: /
    User-agent: Telesoft
    Disallow: /
    User-agent: Website Quester
    Disallow: /
    User-agent: WebZip
    Disallow: /
    User-agent: moget/2.1
    Disallow: /
    User-agent: WebZip/4.0
    Disallow: /
    User-agent: WebStripper
    Disallow: /
    User-agent: WebSauger
    Disallow: /
    User-agent: WebCopier
    Disallow: /
    User-agent: NetAnts
    Disallow: /
    User-agent: Mister PiX
    Disallow: /
    User-agent: WebAuto
    Disallow: /
    User-agent: TheNomad
    Disallow: /
    User-agent: WWW-Collector-E
    Disallow: /
    User-agent: RMA
    Disallow: /
    User-agent: libWeb/clsHTTP
    Disallow: /
    User-agent: asterias
    Disallow: /
    User-agent: httplib
    Disallow: /
    User-agent: turingos
    Disallow: /
    User-agent: spanner
    Disallow: /
    User-agent: InfoNaviRobot
    Disallow: /
    User-agent: Harvest/1.5
    Disallow: /
    User-agent: ExtractorPro
    Disallow: /
    User-agent: Bullseye/1.0
    Disallow: /
    User-agent: Mozilla/4.0 (compatible; BullsEye; Windows 95)
    Disallow: /
    User-agent: Crescent Internet ToolPak HTTP OLE Control v.1.0
    Disallow: /
    User-agent: CherryPickerSE/1.0
    Disallow: /
    User-agent: CherryPickerElite/1.0
    Disallow: /
    User-agent: WebBandit/3.50
    Disallow: /
    User-agent: NICErsPRO
    Disallow: /
    User-agent: Microsoft URL Control - 5.01.4511
    Disallow: /
    User-agent: DittoSpyder
    Disallow: /
    User-agent: Foobot
    Disallow: /
    User-agent: WebmasterWorldForumBot
    Disallow: /
    User-agent: SpankBot
    Disallow: /
    User-agent: BotALot
    Disallow: /
    User-agent: lwp-trivial/1.34
    Disallow: /
    User-agent: lwp-trivial
    Disallow: /
    User-agent: BunnySlippers
    Disallow: /
    User-agent: Microsoft URL Control - 6.00.8169
    Disallow: /
    User-agent: URLy Warning
    Disallow: /
    User-agent: Wget/1.6
    Disallow: /
    User-agent: Wget/1.5.3
    Disallow: /
    User-agent: Wget
    Disallow: /
    User-agent: LinkWalker
    Disallow: /
    User-agent: cosmos
    Disallow: /
    User-agent: moget
    Disallow: /
    User-agent: hloader
    Disallow: /
    User-agent: humanlinks
    Disallow: /
    User-agent: LinkextractorPro
    Disallow: /
    User-agent: Offline Explorer
    Disallow: /
    User-agent: Mata Hari
    Disallow: /
    User-agent: LexiBot
    Disallow: /
    User-agent: Offline Explorer
    Disallow: /
    User-agent: Web Image Collector
    Disallow: /
    User-agent: The Intraformant
    Disallow: /
    User-agent: True_Robot/1.0
    Disallow: /
    User-agent: True_Robot
    Disallow: /
    User-agent: BlowFish/1.0
    Disallow: /
    User-agent: JennyBot
    Disallow: /
    User-agent: MIIxpc/4.2
    Disallow: /
    User-agent: BuiltBotTough
    Disallow: /
    User-agent: ProPowerBot/2.14
    Disallow: /
    User-agent: BackDoorBot/1.0
    Disallow: /
    User-agent: toCrawl/UrlDispatcher
    Disallow: /
    User-agent: WebEnhancer
    Disallow: /
    User-agent: suzuran
    Disallow: /
    User-agent: TightTwatBot
    Disallow: /
    User-agent: VCI WebViewer VCI WebViewer Win32
    Disallow: /
    User-agent: VCI
    Disallow: /
    User-agent: Szukacz/1.4 
    Disallow: /
    User-agent: QueryN Metasearch
    Disallow: /
    User-agent: Openfind data gathere
    Disallow: /
    User-agent: Openfind 
    Disallow: /
    User-agent: Xenu's Link Sleuth 1.1c
    Disallow: /
    User-agent: Xenu's
    Disallow: /
    User-agent: Zeus
    Disallow: /
    User-agent: RepoMonkey Bait & Tackle/v1.01
    Disallow: /
    User-agent: RepoMonkey
    Disallow: /
    User-agent: Microsoft URL Control
    Disallow: /
    User-agent: Openbot
    Disallow: /
    User-agent: URL Control
    Disallow: /
    User-agent: Zeus Link Scout
    Disallow: /
    User-agent: Zeus 32297 Webster Pro V2.9 Win32
    Disallow: /
    User-agent: Webster Pro
    Disallow: /
    User-agent: EroCrawler
    Disallow: /
    User-agent: LinkScan/8.1a Unix
    Disallow: /
    User-agent: Keyword Density/0.9
    Disallow: /
    User-agent: Kenjin Spider
    Disallow: /
    User-agent: Iron33/1.0.2
    Disallow: /
    User-agent: Bookmark search tool
    Disallow: /
    User-agent: GetRight/4.2
    Disallow: /
    User-agent: FairAd Client
    Disallow: /
    User-agent: Gaisbot
    Disallow: /
    User-agent: Aqua_Products
    Disallow: /
    User-agent: Radiation Retriever 1.1
    Disallow: /
  • 26-01-2007, 21:30:23
    #5
    Yazmayı unutmuşum, robots.txt dosyasına bakmayan bazı illegal robotlar engellemen için bu işi .httaccess dosyasından yapman gerekir.

    Ör:

    Options +FollowSymlinks
    RewriteEngine On
    RewriteBase /

    RewriteCond %{HTTP_USER_AGENT} Bot_Referrer_IDsi
    RewriteRule .* - [F,L]

    Bazı robotlar Referrer ID'lerini gizler ya da değiştirirler, bunları engellemek için Robot'un IP'sini ya da IP aralığını bilmen yeterlidir.

    Cyveillance için örnek .httaccess kodu:

    RewriteCond %{REMOTE_ADDR} "^63\.148\.99\.2(2[4-9]|[3-4][0-9]|5[0-5])$"
    RewriteRule .* - [F,L]
  • 26-01-2007, 21:42:58
    #6
    Arkadaşım bildigim kadarı ile google o kadar uzun robots.txt yi sevmiyor. Sen google a izin verir. Geri kalanlarını komple engellersin. ve htacses koruması varsa onu yaparsın. Belki bir nebze indexlemelerine çözüm olur
  • 26-01-2007, 21:50:12
    #7
    Misafir
    saolun arkadaşlar inş işe yarar + sandboxdan çıkabilir domain
  • 02-03-2007, 18:46:33
    #8

    şu siteyi bi incele istersen..