Сообщения

Twitter: 1 billion queries a day and the new engine

At the moment, the load on the servers, Twitter has grown to 1,000 TPS (tweets / sec) and 12,000 QPS (queries per second) — more than 1 billion queries per day. The current infrastructure still stands, but to create a reserve for the next few years, the company has decided to update a backend for a search engine. "If we worked well, you weren't supposed to notice anything in the last week," reported the blog developers Twitter. Until recently, search the backend of Twitter was based on the old SQL system from the company Summize. Her bought in July 2008 for this purpose, and took five of six developers. The need to upgrade Twitter became clear immediately after the presentation of iPhone 3G, and then began a collaboration with Summize. But now it's time to upgrade again. About six months ago, it was decided to develop a new, modern search architecture based on efficient inverted index instead of a relational database. Because Twitter loves open source,...

Index construction for search engines

Full contents and a list of my articles on search engine will be updated here . In previous articles I talked about the work of the search engine here and came to a difficult technical point. Let me remind you that there are 2 types of indexes – direct and reverse. Direct – mapping document a list of words it encountered. Opposite – word is mapped to a list of documents in which it is. Logically, for a quick search of best suited opposite the index. Interesting question and in what order in the list of documents to be stored. In the previous step from the DataFlow module, the indexer, we got a little piece of data in the form of direct index, the reference information and information about the pages. Usually I have it around 200-300mb and contains about 100 thousand pages. Eventually, I abandoned the strategy of storage of whole direct index, and store all of these pieces + full reverse index in several versions so you can revert back. The device index to look at, simpl...

Card Affairs — geographic information service publishing and searching for personal orders

Изображение
Idea, which formed the basis of this service, simple and, I hope, will be in demand by the Internet masses. In short, the essence in the following. This project helps people quickly find assistants, teachers, nurses, promoters, couriers, handymen and other professionals, to assist you in your urgent (or not) businesses for a fee. a few examples Imagine, you need to repair the plumbing, and mechanic of utilities will not wait. What are the alternatives? Look for a locksmith using Yandex/Google? A better way: However, plumbing is not every day change. Example easier You woke up on Saturday and found the last bottle of beer in the fridge... opened it pshshsh... and I realized... well, not you will go in the next couple of hours after yesterday's on the street, to buy something necessary, as in the above the refrigerator mouse hanged. You can suffer, you can: It's simple: open the website www.karta-del.ru , find your house on the map, creat...

I ran the analog Nigma

Изображение
bring to your attention the result of six years of experiments — the search engine "Butty" So was to begin a post with a PR of my engine 2 years ago. A post was already written and ready to publish, but did not work :) Under the cut are waiting for you lyric story of the creation and attempts to promote. Birth the Idea of creating this search engine came into my head in the day of my age. On the same day (2007.08.19) had registered the domain in the zone .EN The establishment of symbiosis search engines at that time already was not new, for example, the search engine Nigma held 0.4% of Russian-speaking search. I called Google and said, whether they admit the combination of issue. The answer was positive. So, for a few months a friend the programmer had written code for a little money. For little money I got unoptimized code, as expected. Further, during the year, we advertise our search engine forums, web-masters, to listen to the opinions ...

AVL trees and the breadth of their application

Decided to describe in my opinion the most useful tree structure. AVL tree is a binary tree (each vertex is not more than 2 sons) in which each vertex has an identifier (and keeps the tree), identifiers obey the following rule: the left son ID<ID of the parent<ID right son. Ie if you traverse a tree recursively from left to right will get sorted in ascending order the list of ID from right to left in descending order. Moreover, the wood is balanced: the height of the left subtree differs from the height right high for 1. I wonder what it means then to check existence of element in the tree takes log(N) N – number ID. It is necessary to go from the root down, and because the tree is maximally symmetric then its height is log(N)+1 The good news is that we're allowed to attach to the top any more useful data and then selection of arbitrary data by ID will take log(N) time. The bad news is the same ID as follows from the definition it can not exist. Will have to...

Spb Transport Online

Изображение
After was launched the city portal of public transport of Saint-Petersburg and installed transport monitoring systems GLONASS/GPS was made possible something that previously could only dream of — coming to a stop to look at where we are right now is the bus that we wait. And "right now" is not just a figure of speech, no exaggeration. Transport really is displayed on the map in real time. Of course, a view of the map on a mobile device. Free software "Spb Transport Online" exists in two versions — for Android and Windows Phone. Despite the different interface, they are very similar in terms of ease-of-use — run, the GPS determines where we are and use the button to select the desired type of transport. The result in the picture above (clickable). Blue, green trolleybuses, and trams — red. example applications the the You are standing at the bus stop at 12 at night waiting for the tram. Are there trams? And if you have to, it is left on the line...

Master class on PostgreSQL developers Skype and other PostgreSQL the October events in Moscow

Изображение
Company Postgresmain" organizing Committee the conference Highload++ are pleased to present to your attention a master class "How to design a scalable architecture PostgreSQL" , which will be conducted by the experts of the company Skype ASKO Oy (Asko Oja) and Marco Creen (Marko Kreen) . The event will be held on 8 October 2008 in Moscow in conference-center "InfoSpace" . "A master class from the developers of Skype will complete a three-day series of PostgreSQL events that we will hold 6-8 October, — says the Executive Director of the company "Postgresmain" Nikolay Samokhvalov . — In the evening of the 6th of October we will organize the next, the fourth open meeting of the Russian community of users of PostgreSQL with the participation of foreign guests. 7th all visitors of the conference Highload++ will be able to listen to an interesting series of reports focused on those using or starting to use PostgreSQL together with our foreig...